AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/5/5-3)
Unit 5 · Topic 5.3
5.3 Linear Regression Models
When a scatterplot looks linear, a regression line ŷ = a + bx turns an x-value into a predicted y-value. This topic covers making predictions, the meaning of ŷ, and why predicting inside the data's x-range (interpolation) is reasonable while predicting outside it (extrapolation) is risky.
Key terms
- regression line ŷ = a + bx
- predicted value (ŷ)
- slope (b) and y-intercept (a)
- interpolation
- extrapolation
The regression model
A linear regression model is ŷ = a + bx. Here x is the explanatory variable, ŷ ("y-hat") is the predicted response, a is the y-intercept and b is the slope. The hat is important: ŷ is a prediction, not an actual observed value.
Only fit a line when the scatterplot looks linear. A line through curved data gives predictions that are systematically wrong in places.
Statistics books usually write the intercept first (a + bx). Some calculators and software write it as y = ax + b or show a table of coefficients. Read the labels carefully to tell which number is the slope.
Making predictions
Plug the x-value into the equation. For the hours-and-scores data, the least-squares line is ŷ = 58.55 + 3.83x. For a student who studies 4.5 hours, ŷ = 58.55 + 3.83(4.5) ≈ 75.8 points.
Write it in context: "The predicted exam score for a student who studies 4.5 hours is about 75.8 points." Don't round in a way that changes the meaning, and include units.
Interpolation vs. extrapolation
Interpolation is predicting for an x-value inside the range of x-values used to build the line. The data in this example run from 1 to 8 hours, so predicting at 4.5 hours is interpolation.
Extrapolation is predicting for an x-value outside that range. Predicting a score for 20 hours of study gives ŷ = 58.55 + 3.83(20) ≈ 135, which is impossible on a 100-point exam. The linear pattern may not continue beyond the data you have, and the farther you go, the less reliable the prediction.
Predictions are averages
ŷ estimates the typical response for individuals with that x-value. Two students who each studied 5 hours scored 74 and 78, and the line predicts 77.7 for both. Individuals vary around the line; the gap is a residual (topic 5.4).
Before you predict
Run through this checklist before you trust a prediction. Predictions also only apply to individuals like the ones in the data. A line built from high school students says little about college students or professional test-takers.
- Check the scatterplot: is the form linear?
- Check the residual plot (5.4): no curved pattern?
- Check the x-value: is it inside the range of the data?
- Check the strength: a weak association gives predictions with lots of error, even inside the range.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Predict and judge reliability
The regression line for exam score on hours studied (data from 1 to 8 hours) is ŷ = 58.55 + 3.83x. Predict the score for a student who studies 6.5 hours and for one who studies 12 hours. Which prediction is more trustworthy?
Show the solutionHide the solution
- Step 1: 6.5 hours: ŷ = 58.55 + 3.83(6.5) ≈ 83.4 points. 6.5 is inside 1–8 hours, so this is interpolation.
- Step 2: 12 hours: ŷ = 58.55 + 3.83(12) ≈ 104.5 points. 12 is outside the data, so this is extrapolation, and the result is impossible on a 100-point exam.
- Step 3: The 6.5-hour prediction is reasonable; the 12-hour prediction isn't.
Answer: About 83.4 points at 6.5 hours (interpolation, reasonable) and about 104.5 at 12 hours (extrapolation, unreliable and impossible).
- Example 2Calculator allowed
Trap: reading the slope as the intercept
Software output for a regression of weight (kg) on height (cm) shows Constant = −102.4 and Height = 0.98. A student writes ŷ = 0.98 − 102.4x. Write the correct equation and predict the weight of someone 175 cm tall.
Show the solutionHide the solution
- Step 1: "Constant" is the y-intercept a; the coefficient next to the variable name is the slope b.
- Step 2: Correct equation: predicted weight = −102.4 + 0.98(height).
- Step 3: At 175 cm: −102.4 + 0.98(175) = −102.4 + 171.5 = 69.1 kg.
Answer: ŷ = −102.4 + 0.98x; predicted weight ≈ 69.1 kg.
Common mistakes
- Writing the equation without the hat (y instead of ŷ) or without defining x and y in context.
- Mixing up the slope and the intercept in software output.
- Trusting predictions far outside the range of the data.
- Treating ŷ as the exact value every individual with that x will have.
On the exam
- When you write a regression equation, define the variables: "predicted exam score = 58.55 + 3.83(hours studied)."
- If a question asks whether a prediction is reasonable, check whether the x-value is inside the data's range and say "extrapolation" when it isn't.
Connected topics
Videos
Check yourself
4 questions on 5.3 Linear Regression Models. Pick an answer to see if you got it, and why.
A regression line for predicting a tree's height (feet) from its trunk diameter (inches) is ŷ = 4.2 + 2.6x, based on trees with diameters from 4 to 20 inches. What is the predicted height of a tree with a 12-inch diameter?
Using the same tree line, ŷ = 4.2 + 2.6x, for diameters from 4 to 20 inches, which prediction is extrapolation?
For the regression line ŷ = 120 − 3.5x, where x is the number of days since a pond was treated and ŷ is the predicted algae count per sample, what are the slope and the y-intercept?
A regression of a child's height on age, using children from 2 to 10 years old, gives a good linear fit. A parent uses the line to predict her child's height at age 40. Why is this a poor idea?
0 of 4 answered