Skip to main content

Unit 5 · Topic 5.1

5.1 Graphical Representations Between Two Quantitative Variables

A scatterplot shows how two quantitative variables measured on the same individuals relate. You'll learn which variable goes on which axis and how to describe an association by its form, direction, strength and unusual features, always in context.

Key terms

  • scatterplot
  • explanatory variable
  • response variable
  • form, direction and strength
  • unusual features

Bivariate data and scatterplots

Bivariate quantitative data are pairs of numbers measured on the same individuals, like each student's hours studied and exam score. Each individual becomes one point (x, y) on a scatterplot.

The explanatory variable goes on the horizontal x-axis. It's the variable you use to explain or predict the other one. The response variable goes on the vertical y-axis. If there's no clear explanatory variable (arm span and height, say), either can go on either axis.

Label both axes with the variable names and units, and use scales that start near the smallest values instead of forcing them to 0 when that would squash the data.

Hours studied (x)1223455678
Exam score (y)62657068757874838590

Describing an association

Cover four things, in context:

  • Form: linear (points scatter around a straight line) or nonlinear (they follow a curve).
  • Direction: positive (as x increases, y tends to increase) or negative (as x increases, y tends to decrease). Some curved patterns have no single direction.
  • Strength: how closely the points follow the pattern: strong, moderate or weak.
  • Unusual features: clusters, or individual points that don't fit the overall pattern.

Say it in context

"There is a strong, positive, linear association between hours studied and exam score for these 10 students, with no unusual points" beats "strong positive linear."

Even better, explain the direction in plain words: "students who studied more hours tended to earn higher scores." Words like "tend to" matter. A positive association doesn't mean every student who studied more scored higher; it's a general trend.

Unusual features

Look for two kinds. A cluster is a group of points separated from the rest, which can mean the data mix two different kinds of individuals (say, cars and trucks in a plot of weight vs. fuel efficiency). A point that doesn't fit is one far from the overall pattern, like a student who studied 7 hours but scored 55.

When you mention an unusual point, describe it in context and say how it departs from the pattern: "one student studied 7 hours but scored much lower than students with similar study time." Don't delete it; ask whether it's an error or a real, interesting case.

What a scatterplot can't tell you

A scatterplot shows association, not causation. If the data come from an observational study, a third variable could drive both. Ice cream sales and swimming-pool accidents both rise in hot weather, so they're associated even though ice cream doesn't cause accidents.

Don't call a relationship linear just because you could draw a line through it. Look for whether the points curve away from a line at either end.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Describe a scatterplot

    Using the hours-and-scores data above, describe the scatterplot of exam score (y) vs. hours studied (x).

    Show the solution
    1. Step 1: Hours studied explains score, so it goes on the x-axis.
    2. Step 2: Plotting the 10 points shows them rising from lower left (1 hour, 62) to upper right (8 hours, 90), close to a straight line.
    3. Step 3: Form: linear. Direction: positive. Strength: strong, since points stay close to the line. Unusual features: none; no point stands apart.

    Answer: There is a strong, positive, linear association between hours studied and exam score for these students: those who studied more tended to score higher. No unusual points stand out.

  2. Example 2Calculator allowed

    Trap: describing a curve as linear

    A scatterplot of a car's fuel efficiency (mpg) vs. speed (mph) rises from 20 mph to a peak near 50 mph, then falls as speed increases to 80 mph. A student describes it as "a weak, positive, linear association." Correct the description.

    Show the solution
    1. Step 1: The points follow a curve that goes up and then down, so the form is nonlinear.
    2. Step 2: There's no single direction: efficiency increases with speed up to about 50 mph, then decreases.
    3. Step 3: The pattern itself can be strong even though a straight line would fit badly.

    Answer: The association between speed and fuel efficiency is nonlinear: efficiency rises until about 50 mph and then falls. It isn't well described as linear or as simply positive.

Common mistakes

  • Putting the response variable on the x-axis.
  • Describing a scatterplot without context ("it's positive").
  • Calling a curved pattern weak because a line doesn't fit it.
  • Claiming one variable causes the other from an observational scatterplot.

On the exam

  • "Describe the relationship" means form, direction, strength and unusual features, in context. Writing "positive" alone usually doesn't earn credit; say what increases with what.
  • Regression questions often come as a set of connected multiple-choice questions built on one data set, so read the context carefully the first time.

Connected topics

Videos

  • AP Statistics Topic 5.1 Graphical Representations Between Two Quantitative Variables | Scatterplots

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • AP Stats 5.1 - Describing Two Quantitative Variables

    Skew The ScriptWatch on YouTube (opens in a new tab)

  • Bivariate relationship linearity, strength and direction | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • Explanatory and Response Variables, Correlation (2.1)

    Simple Learning ProWatch on YouTube (opens in a new tab)

  • Describing Scatterplots [AP Statistics Topic 2.4]

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • Constructing a scatter plot | Regression | Probability and Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 5.1 Graphical Representations Between Two Quantitative Variables. Pick an answer to see if you got it, and why.

Question 1 of 4Calculator allowed

A researcher studies whether the amount of rainfall during the growing season helps predict corn yield on farms. Which is the correct way to set up the scatterplot?

Question 2 of 4Calculator allowed

A scatterplot of outside temperature (°F) and a home's daily heating cost ($) for 60 winter days shows points falling from upper left to lower right in a fairly tight straight-line band, with no unusual points. Which is the best description?

Question 3 of 4Calculator allowed

A scatterplot of the weights and fuel economy of 40 vehicles shows a moderate negative linear pattern. Eight of the vehicles, all electric, form a separate group in the upper right with very high fuel-economy ratings. How should this feature be described?

Question 4 of 4Calculator allowed

For a random sample of adults, which pair of variables would most likely show a negative association?

0 of 4 answered