AP® Computer Science A review sheet from Aim for Five (aimforfive.com/csa/units/4/4-1)
Unit 4 · Topic 4.1
4.1 Ethical and Social Issues Around Data Collection
Programs that collect and use data affect real people. This topic covers the privacy risks of storing personal data, how bias and bad data creep into programs, and how to judge whether a data set actually fits the question you're trying to answer.
Key terms
- privacy
- personal data
- algorithmic bias
- data quality
Privacy
Every time someone uses an app or a website, some of their personal information can be exposed. Apps can record names, locations, contacts, searches and habits, and once data is stored, it can be leaked, sold, stolen or combined with other data to learn far more than the user expected.
Programmers should try to protect the people who use their programs. Practical habits include collecting only the data the program really needs, keeping it secure, deleting it when it's no longer needed, and being clear with users about what's collected and why.
Algorithmic bias
Algorithmic bias is when a program keeps getting things wrong in the same way for one group of people, so that group is treated unfairly. It isn't one random mistake; it's a pattern.
Bias often starts with the data. If a face-recognition program is trained mostly on photos of one group of people, it may work well for them and poorly for everyone else. If a hiring program learns from a company's past hires, it can copy whatever unfairness was in those past decisions.
So before using data to draw conclusions, ask how it was collected. An online survey only hears from people who are online and choose to answer. Sensor data only covers places that have sensors.
Data quality
Some data sets are incomplete or contain mistakes: missing values, typos, duplicates, or impossible numbers like a height of 0. A program that uses bad data can give wrong answers or run inefficiently, even when its code is perfect.
Missing values are often stored as a placeholder like -1 or 0. If your code doesn't skip them, they get averaged in as if they were real.
Data can also be incomplete in a bigger way. A survey about students' sleep taken only during exam week, or a fitness study that only includes people who already own a smartwatch, misses part of the picture. The code may run perfectly and still produce conclusions that don't hold for everyone.
Choosing the right data set
A data set is gathered for a purpose. Data that answers one question well may be useless, or misleading, for another. A set of school cafeteria sales can tell you which lunches are popular, but not what students eat at home or what they'd buy if prices changed.
Worked examples
Try each one yourself first, then open the solution.
- Example 1
Placeholder values distort the answer
A study app records minutes studied each day, using -1 for days when the student forgot to log. What does this code print, and which number is the honest average?
int[] minutes = {32, 45, -1, 28, -1, 50}; int sum = 0; for (int m : minutes) { sum += m; } System.out.println((double) sum / minutes.length); int realSum = 0; int count = 0; for (int m : minutes) { if (m != -1) { realSum += m; count++; } } System.out.println((double) realSum / count);Show the solutionHide the solution
- Step 1: The first loop adds every value, including the two -1s: 32 + 45 − 1 + 28 − 1 + 50 = 153. Dividing by all 6 entries gives 25.5.
- Step 2: The second loop skips the -1s. The real values are 32, 45, 28 and 50, which add to 155. There are 4 of them, so the average is 155 / 4 = 38.75.
- Step 3: The -1s aren't study times, they're missing data. Counting them pulls the average down and also adds two fake days to the count.
Answer: It prints
25.5and then38.75. The honest average is 38.75 minutes, over the 4 days that were actually logged. - Example 2
Is this data set right for the question?
A city wants to know how many residents would ride a new bus route. It plans to use data from its bike-share app, which records every ride taken. Explain one problem with using this data set.
Show the solutionHide the solution
- Step 1: Ask who's in the data. The bike-share data only includes people who use the bike-share app.
- Step 2: Many likely bus riders, such as older residents, people with disabilities, or people without a smartphone or credit card, may never appear in it.
- Step 3: The data was collected to answer a different question (how people use bikes), so using it to predict bus riders could give a biased, misleading estimate.
Answer: The data set leaves out many people who might ride the bus, because it only records bike-share users. It was gathered for a different question, so conclusions about bus riders could be biased.
Common mistakes
- Thinking bias only comes from biased programmers. It often comes from the data itself or from how the data was collected.
- Treating placeholder values like -1 as real data in totals, averages or minimums.
- Assuming any big data set can answer any question. Check that it was collected for the question you're asking.
On the exam
- Multiple-choice questions may describe a program or data set and ask which is a privacy risk, an example of bias, or a reason the data is unsuitable. Look for who is left out and what the data was collected for.
Connected topics
Videos
Check yourself
4 questions on 4.1 Ethical and Social Issues Around Data Collection. Pick an answer to see if you got it, and why.
A company uses a program to rank job applicants. The program learned from the company's past hiring decisions, and it now gives consistently lower rankings to qualified applicants from one neighborhood. Which of the following best describes this problem?
A team is building a fitness app that records users' locations during runs. Which of the following practices best protects users' personal privacy?
A student has a data set from a survey of 200 high school seniors in one city about their favorite snacks. She wants to use it to predict which snacks adults across the whole country prefer. Which of the following is the biggest problem with her plan?
A weather station records one temperature per hour. When the sensor fails, it records -999 instead of a temperature. A program computes the day's average temperature by adding all 24 values and dividing by 24. Which of the following best describes the result on a day when the sensor failed twice?
0 of 4 answered