AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/4/4-10)
Unit 4 · Topic 4.10
4.10 Carrying Out a Test for the Difference Between Two Population Means
To carry out a two-sample t-test, compute t = (x̄₁ − x̄₂ − 0)/√(s₁²/n₁ + s₂²/n₂), get the p-value from a t-distribution (usually with technology), compare it with α and write a conclusion in context about the two population means.
Key terms
- t test statistic
- p-value
- significance level (α)
- conclusion in context
The test statistic
t = (x̄₁ − x̄₂ − 0) ÷ √(s₁²/n₁ + s₂²/n₂). The 0 is the hypothesized difference from H₀. The denominator is the same standard error used in the interval; for means, there's no pooling.
When H₀ is true, t has approximately a t-distribution. Technology (2-SampTTest, with Pooled: No) gives the df, which falls between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2. By hand, using the smaller of n₁ − 1 and n₂ − 1 is the safe choice.
p-value and decision
The p-value is the area beyond t in the direction of Hₐ. Interpret it with the null assumption: "Assuming the true mean commute times in the two cities are equal, there is about a 0.016 probability of getting a difference in sample means of 4.3 minutes or more by chance alone."
If the p-value ≤ α, reject H₀: there is convincing evidence for Hₐ. Otherwise, fail to reject: there isn't convincing evidence for Hₐ.
Conclusion and scope
Write the conclusion about the population means, in context, matching Hₐ: "There is convincing evidence that the true mean commute time for all workers in City A is greater than for all workers in City B."
With two random samples, you can generalize to the two populations but can't claim the city causes longer commutes. With a randomized experiment, rejecting H₀ supports cause and effect for units like those in the study.
Errors and power
Rejecting H₀ risks a Type I error (concluding the means differ when they don't). Failing to reject risks a Type II error (missing a real difference). Larger samples, less variable data, a larger true difference or a larger α all increase power, just as in topic 3.8.
Matching the interval
A two-sided test at α = 0.05 and a 95% two-sample t-interval usually agree: the test rejects when 0 is outside the interval. A one-sided test at α = 0.05 matches a 90% interval instead, because it puts all 5% in one tail.
Reading calculator output
2-SampTTest reports t, the p-value, df, both sample means and both standard deviations. On free response, copy the important numbers into a sentence and show the formula; "2-SampTTest gave p = 0.016" alone usually doesn't earn the mechanics point.
Make sure the order you entered the groups matches your hypotheses. Entering City B first would flip the sign of t and the direction of the test.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Carry out a two-sample t-test
Continue the commute test from 4.9. City A: n = 35, x̄ = 28.4 min, s = 9.2 min. City B: n = 40, x̄ = 24.1 min, s = 7.5 min. Test H₀: μA − μB = 0 vs. Hₐ: μA − μB > 0 at α = 0.05.
Show the solutionHide the solution
- Step 1: SE = √(9.2²/35 + 7.5²/40) ≈ 1.956.
- Step 2: t = (28.4 − 24.1 − 0)/1.956 ≈ 2.20.
- Step 3: Technology: df ≈ 65.7, p-value ≈ 0.016. (With the conservative df = 34, p ≈ 0.017.)
- Step 4: 0.016 < 0.05, so reject H₀.
Answer: t ≈ 2.20, df ≈ 65.7, p-value ≈ 0.016. There is convincing evidence that the true mean commute time for all workers in City A is greater than for all workers in City B.
- Example 2Calculator allowed
Trap: claiming causation from samples
After the commute test, a student writes: "Living in City A causes people to have longer commutes." Evaluate this.
Show the solutionHide the solution
- Step 1: The data come from random samples, not an experiment; nobody was assigned to a city.
- Step 2: Other variables could explain the difference, such as city size, traffic or job locations.
- Step 3: The test supports a difference in population means, not a cause.
Answer: Not justified. Random sampling lets you generalize the difference to both cities' workers, but without random assignment you can't conclude the city causes longer commutes.
Common mistakes
- Using a pooled standard error or the pooled calculator option.
- Using df = n₁ + n₂ − 2 by hand.
- Writing the conclusion about sample means or about individuals rather than population means.
- Claiming causation without random assignment.
On the exam
- A complete answer shows the test statistic formula with values, t, df and the p-value, then a conclusion linked to α in context.
- Inference free-response questions often add a part about scope or errors. Be ready to say whether cause and effect or generalization is justified.
Connected topics
- Unit 33.13 Carrying Out a Test for the Difference Between Two Population Proportions
- Unit 44.5 Carrying Out a Test for a Population Mean or Population Mean Difference
- Unit 44.8 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means
- Unit 44.9 Setting Up a Test for the Difference Between Two Population Means
Videos
Check yourself
4 questions on 4.10 Carrying Out a Test for the Difference Between Two Population Means. Pick an answer to see if you got it, and why.
A randomized experiment compares mean plant growth with a new fertilizer and with a standard fertilizer. A one-sided test of H₀: μ_new = μ_standard versus Hₐ: μ_new > μ_standard gives p-value = 0.008. Which conclusion is correct at α = 0.05?
Independent random samples of students who do and don't play a musical instrument show a mean GPA that is significantly higher for instrument players (p-value = 0.002). Which conclusion is appropriate?
A two-sample t-test comparing two means gives a p-value of 0.09. A researcher repeats the study with much larger samples, and the sample means and standard deviations come out about the same as before. What is most likely to happen to the p-value?
A two-sample t-test of H₀: μ₁ = μ₂ versus Hₐ: μ₁ > μ₂ gives a p-value of 0.03. Which is the correct interpretation?
0 of 4 answered