Skip to main content

Unit 4 · Topic 4.1

4.1 Ethical and Social Issues Around Data Collection

Programs that collect and use data affect real people. This topic covers the privacy risks of storing personal data, how bias and bad data creep into programs, and how to judge whether a data set actually fits the question you're trying to answer.

Key terms

  • privacy
  • personal data
  • algorithmic bias
  • data quality

Privacy

Every time someone uses an app or a website, some of their personal information can be exposed. Apps can record names, locations, contacts, searches and habits, and once data is stored, it can be leaked, sold, stolen or combined with other data to learn far more than the user expected.

Programmers should try to protect the people who use their programs. Practical habits include collecting only the data the program really needs, keeping it secure, deleting it when it's no longer needed, and being clear with users about what's collected and why.

Algorithmic bias

Algorithmic bias is when a program keeps getting things wrong in the same way for one group of people, so that group is treated unfairly. It isn't one random mistake; it's a pattern.

Bias often starts with the data. If a face-recognition program is trained mostly on photos of one group of people, it may work well for them and poorly for everyone else. If a hiring program learns from a company's past hires, it can copy whatever unfairness was in those past decisions.

So before using data to draw conclusions, ask how it was collected. An online survey only hears from people who are online and choose to answer. Sensor data only covers places that have sensors.

Data quality

Some data sets are incomplete or contain mistakes: missing values, typos, duplicates, or impossible numbers like a height of 0. A program that uses bad data can give wrong answers or run inefficiently, even when its code is perfect.

Missing values are often stored as a placeholder like -1 or 0. If your code doesn't skip them, they get averaged in as if they were real.

Data can also be incomplete in a bigger way. A survey about students' sleep taken only during exam week, or a fitness study that only includes people who already own a smartwatch, misses part of the picture. The code may run perfectly and still produce conclusions that don't hold for everyone.

Choosing the right data set

A data set is gathered for a purpose. Data that answers one question well may be useless, or misleading, for another. A set of school cafeteria sales can tell you which lunches are popular, but not what students eat at home or what they'd buy if prices changed.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1

    Placeholder values distort the answer

    A study app records minutes studied each day, using -1 for days when the student forgot to log. What does this code print, and which number is the honest average?int[] minutes = {32, 45, -1, 28, -1, 50}; int sum = 0; for (int m : minutes) { sum += m; } System.out.println((double) sum / minutes.length); int realSum = 0; int count = 0; for (int m : minutes) { if (m != -1) { realSum += m; count++; } } System.out.println((double) realSum / count);

    Show the solution
    1. Step 1: The first loop adds every value, including the two -1s: 32 + 45 − 1 + 28 − 1 + 50 = 153. Dividing by all 6 entries gives 25.5.
    2. Step 2: The second loop skips the -1s. The real values are 32, 45, 28 and 50, which add to 155. There are 4 of them, so the average is 155 / 4 = 38.75.
    3. Step 3: The -1s aren't study times, they're missing data. Counting them pulls the average down and also adds two fake days to the count.

    Answer: It prints 25.5 and then 38.75. The honest average is 38.75 minutes, over the 4 days that were actually logged.

  2. Example 2

    Is this data set right for the question?

    A city wants to know how many residents would ride a new bus route. It plans to use data from its bike-share app, which records every ride taken. Explain one problem with using this data set.

    Show the solution
    1. Step 1: Ask who's in the data. The bike-share data only includes people who use the bike-share app.
    2. Step 2: Many likely bus riders, such as older residents, people with disabilities, or people without a smartphone or credit card, may never appear in it.
    3. Step 3: The data was collected to answer a different question (how people use bikes), so using it to predict bus riders could give a biased, misleading estimate.

    Answer: The data set leaves out many people who might ride the bus, because it only records bike-share users. It was gathered for a different question, so conclusions about bus riders could be biased.

Common mistakes

  • Thinking bias only comes from biased programmers. It often comes from the data itself or from how the data was collected.
  • Treating placeholder values like -1 as real data in totals, averages or minimums.
  • Assuming any big data set can answer any question. Check that it was collected for the question you're asking.

On the exam

  • Multiple-choice questions may describe a program or data set and ask which is a privacy risk, an example of bias, or a reason the data is unsuitable. Look for who is left out and what the data was collected for.

Connected topics

Videos

  • AP Computer Science A - Topic 4.1: Ethical and Social Issues Around Data Collection

    Tim Gallagher Computer ScienceWatch on YouTube (opens in a new tab)

  • AP CSA Data Collections – Ethics and Social Issues Around Data Collection

    Goldie's Math EmporiumWatch on YouTube (opens in a new tab)

  • Algorithmic Bias and Fairness: Crash Course AI #18

    CrashCourseWatch on YouTube (opens in a new tab)

  • AP CS A - 7.7 Ethical Issues Around Data Collection

    CodeHSWatch on YouTube (opens in a new tab)

  • Algorithmic bias | Intro to CS - Python | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 4.1 Ethical and Social Issues Around Data Collection. Pick an answer to see if you got it, and why.

Question 1 of 4

A company uses a program to rank job applicants. The program learned from the company's past hiring decisions, and it now gives consistently lower rankings to qualified applicants from one neighborhood. Which of the following best describes this problem?

Question 2 of 4

A team is building a fitness app that records users' locations during runs. Which of the following practices best protects users' personal privacy?

Question 3 of 4

A student has a data set from a survey of 200 high school seniors in one city about their favorite snacks. She wants to use it to predict which snacks adults across the whole country prefer. Which of the following is the biggest problem with her plan?

Question 4 of 4

A weather station records one temperature per hour. When the sensor fails, it records -999 instead of a temperature. A program computes the day's average temperature by adding all 24 values and dividing by 24. Which of the following best describes the result on a day when the sensor failed twice?

0 of 4 answered