AP® Computer Science Principles review sheet from Aim for Five (aimforfive.com/csp/units/2/2-3)
Unit 2 · Topic 2.3
2.3 Extracting Information from Data
Data becomes useful once you pull information out of it: facts, trends and patterns. This topic covers metadata, the work of cleaning and combining data, the limits of what data can tell you (bias and correlation versus causation), and why large data sets need bigger systems.
Key terms
- information
- metadata
- data cleaning
- correlation vs. causation
- bias in data
- scalability
From data to information
Data are raw values: numbers, text, readings, clicks. Information is the facts and patterns you extract from data. A list of 10,000 bus tap-in times is data; "the 7:40 bus is overcrowded on Mondays" is information.
Data can reveal trends, show connections between things and help address real problems, like where to add a crosswalk or when a store needs more staff.
Often one data source isn't enough. You may need to combine sources to draw a conclusion. To decide whether a park needs more lighting, a city might combine crime reports with streetlight locations and park-visit counts.
Metadata
Metadata are data about data. For a photo, the data is the image itself; the metadata might include the date and time it was taken, the GPS location, the camera model and the file size. For a song, metadata might be the title, artist, length and genre.
Metadata help you find, organize and manage data. They're how a phone groups photos by place or a music app sorts by artist. They also add context that makes data more useful.
Changing or deleting metadata doesn't change the data itself. Editing a photo's date tag doesn't change a single pixel.
Cleaning and other challenges
Every data set, large or small, can have problems:
- Data that needs cleaning: when people type into an open text box, the same answer shows up as "NY," "N.Y." and "new york." Cleaning makes data uniform without changing its meaning, such as by replacing all versions with one spelling.
- Incomplete data: missing values, like survey questions people skipped.
- Invalid data: values that can't be right, like an age of 250 or a date of February 31.
- Combining sources: different sources may use different formats or units that have to be matched up first.
Limits: bias and correlation
Bias often comes from the type or source of the data collected. If a survey about school lunches only reaches students who eat in the cafeteria, it leaves out everyone who doesn't, and collecting more responses the same way won't fix that. More data doesn't remove bias.
Data may show a correlation, meaning two things tend to change together. That doesn't prove one causes the other. Ice cream sales and sunburns rise together, but both are caused by hot, sunny weather. You'd need more research to know the real relationship.
Size and scalability
The size of a data set affects how much information you can get from it. Bigger data sets can reveal patterns a small sample would miss.
Very large data sets are hard to process on one computer and may need parallel systems, where many processors or computers share the work. Scalability is a system's ability to grow to handle more data or users. How much a system can compute and store limits what you can do with a data set.
What you can do with data also depends on the people and tools available. A spreadsheet works for a few thousand rows; millions of rows may need a database and someone who can program it.
Worked examples
Try each one yourself first, then open the solution.
- Example 1
What can the data tell you?
A photo-sharing app stores each uploaded photo along with this metadata: the date and time, the GPS location, and the user's ID. Which of these questions can be answered using only the metadata? (1) Which city had the most uploads last summer? (2) How many photos show a dog? (3) Which users upload photos most often late at night?
Show the solutionHide the solution
- Step 1: Question 1 needs locations and dates. Both are in the metadata, so you can count uploads per city during the summer months.
- Step 2: Question 2 needs to know what's in each picture. That's in the image data, not the metadata, so the metadata alone can't answer it.
- Step 3: Question 3 needs user IDs and times. Both are in the metadata, so you can count each user's uploads in late-night hours.
Answer: Questions 1 and 3 can be answered from the metadata; question 2 cannot, because it depends on the image content.
Common mistakes
- Concluding that one thing causes another because they're correlated. Correlation only shows they change together.
- Thinking a bigger data set removes bias. If the collection method leaves people out, more of the same data repeats the bias.
- Saying that editing metadata changes the original data. They're stored separately.
- Treating cleaning as changing the data's meaning. Cleaning only makes equivalent entries consistent.
On the exam
- A very common question type gives a data set and its metadata, then asks which question it can (or can't) answer. Check whether each needed piece of information is actually in the data described.
- Expect questions about why a conclusion is unreliable: watch for biased collection, missing data, or a claim of cause based only on correlation.
Connected topics
Videos
Check yourself
4 questions on 2.3 Extracting Information from Data. Pick an answer to see if you got it, and why.
A digital photo is stored along with the following: the date and time it was taken, the GPS location, the camera model and the file size. These items are best described as
A user edits a song file's metadata to correct the artist's name. What effect does this have on the song's audio data?
A survey let students type their grade level. The responses include "9", "9th", "ninth" and "Freshman." Before counting students by grade, a programmer changes all of these to "9." This process is called
A city's data shows that in months when more ice cream is sold, more people visit the emergency room for sunburn. Which conclusion is best supported by this data alone?
0 of 4 answered