Skip to main content

Unit 5 · Topic 5.3

5.3 Linear Regression Models

When a scatterplot looks linear, a regression line ŷ = a + bx turns an x-value into a predicted y-value. This topic covers making predictions, the meaning of ŷ, and why predicting inside the data's x-range (interpolation) is reasonable while predicting outside it (extrapolation) is risky.

Key terms

  • regression line ŷ = a + bx
  • predicted value (ŷ)
  • slope (b) and y-intercept (a)
  • interpolation
  • extrapolation

The regression model

A linear regression model is ŷ = a + bx. Here x is the explanatory variable, ŷ ("y-hat") is the predicted response, a is the y-intercept and b is the slope. The hat is important: ŷ is a prediction, not an actual observed value.

Only fit a line when the scatterplot looks linear. A line through curved data gives predictions that are systematically wrong in places.

Statistics books usually write the intercept first (a + bx). Some calculators and software write it as y = ax + b or show a table of coefficients. Read the labels carefully to tell which number is the slope.

Making predictions

Plug the x-value into the equation. For the hours-and-scores data, the least-squares line is ŷ = 58.55 + 3.83x. For a student who studies 4.5 hours, ŷ = 58.55 + 3.83(4.5) ≈ 75.8 points.

Write it in context: "The predicted exam score for a student who studies 4.5 hours is about 75.8 points." Don't round in a way that changes the meaning, and include units.

Interpolation vs. extrapolation

Interpolation is predicting for an x-value inside the range of x-values used to build the line. The data in this example run from 1 to 8 hours, so predicting at 4.5 hours is interpolation.

Extrapolation is predicting for an x-value outside that range. Predicting a score for 20 hours of study gives ŷ = 58.55 + 3.83(20) ≈ 135, which is impossible on a 100-point exam. The linear pattern may not continue beyond the data you have, and the farther you go, the less reliable the prediction.

Predictions are averages

ŷ estimates the typical response for individuals with that x-value. Two students who each studied 5 hours scored 74 and 78, and the line predicts 77.7 for both. Individuals vary around the line; the gap is a residual (topic 5.4).

Before you predict

Run through this checklist before you trust a prediction. Predictions also only apply to individuals like the ones in the data. A line built from high school students says little about college students or professional test-takers.

  • Check the scatterplot: is the form linear?
  • Check the residual plot (5.4): no curved pattern?
  • Check the x-value: is it inside the range of the data?
  • Check the strength: a weak association gives predictions with lots of error, even inside the range.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Predict and judge reliability

    The regression line for exam score on hours studied (data from 1 to 8 hours) is ŷ = 58.55 + 3.83x. Predict the score for a student who studies 6.5 hours and for one who studies 12 hours. Which prediction is more trustworthy?

    Show the solution
    1. Step 1: 6.5 hours: ŷ = 58.55 + 3.83(6.5) ≈ 83.4 points. 6.5 is inside 1–8 hours, so this is interpolation.
    2. Step 2: 12 hours: ŷ = 58.55 + 3.83(12) ≈ 104.5 points. 12 is outside the data, so this is extrapolation, and the result is impossible on a 100-point exam.
    3. Step 3: The 6.5-hour prediction is reasonable; the 12-hour prediction isn't.

    Answer: About 83.4 points at 6.5 hours (interpolation, reasonable) and about 104.5 at 12 hours (extrapolation, unreliable and impossible).

  2. Example 2Calculator allowed

    Trap: reading the slope as the intercept

    Software output for a regression of weight (kg) on height (cm) shows Constant = −102.4 and Height = 0.98. A student writes ŷ = 0.98 − 102.4x. Write the correct equation and predict the weight of someone 175 cm tall.

    Show the solution
    1. Step 1: "Constant" is the y-intercept a; the coefficient next to the variable name is the slope b.
    2. Step 2: Correct equation: predicted weight = −102.4 + 0.98(height).
    3. Step 3: At 175 cm: −102.4 + 0.98(175) = −102.4 + 171.5 = 69.1 kg.

    Answer: ŷ = −102.4 + 0.98x; predicted weight ≈ 69.1 kg.

Common mistakes

  • Writing the equation without the hat (y instead of ŷ) or without defining x and y in context.
  • Mixing up the slope and the intercept in software output.
  • Trusting predictions far outside the range of the data.
  • Treating ŷ as the exact value every individual with that x will have.

On the exam

  • When you write a regression equation, define the variables: "predicted exam score = 58.55 + 3.83(hours studied)."
  • If a question asks whether a prediction is reasonable, check whether the x-value is inside the data's range and say "extrapolation" when it isn't.

Connected topics

Videos

  • AP Statistics Topic 5.3 Linear Regression Models | Lesson + Guided Notes + Practice Problems

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • AP Stats 5.2 - Linear Regression & Residuals

    Skew The ScriptWatch on YouTube (opens in a new tab)

  • Linear Regression Models Explained for AP Statistics Topic 2.6

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • Introduction to Simple Linear Regression

    jbstatisticsWatch on YouTube (opens in a new tab)

  • Example estimating from regression line

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 5.3 Linear Regression Models. Pick an answer to see if you got it, and why.

Question 1 of 4Calculator allowed

A regression line for predicting a tree's height (feet) from its trunk diameter (inches) is ŷ = 4.2 + 2.6x, based on trees with diameters from 4 to 20 inches. What is the predicted height of a tree with a 12-inch diameter?

Question 2 of 4Calculator allowed

Using the same tree line, ŷ = 4.2 + 2.6x, for diameters from 4 to 20 inches, which prediction is extrapolation?

Question 3 of 4Calculator allowed

For the regression line ŷ = 120 − 3.5x, where x is the number of days since a pond was treated and ŷ is the predicted algae count per sample, what are the slope and the y-intercept?

Question 4 of 4Calculator allowed

A regression of a child's height on age, using children from 2 to 10 years old, gives a good linear fit. A parent uses the line to predict her child's height at age 40. Why is this a poor idea?

0 of 4 answered