AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/4/4-6)
Unit 4 · Topic 4.6
4.6 Sampling Distributions for the Difference Between Two Sample Means
To compare two groups on a quantitative variable, look at the difference in sample means, x̄₁ − x̄₂. Its sampling distribution is centered at μ₁ − μ₂, has standard deviation √(σ₁²/n₁ + σ₂²/n₂), and is normal or approximately normal under the right conditions.
Key terms
- difference in sample means (x̄₁ − x̄₂)
- independent samples
- standard deviation √(σ₁²/n₁ + σ₂²/n₂)
- normality condition
Center and spread
For two independent random samples (or two groups formed by random assignment), the sampling distribution of x̄₁ − x̄₂ has mean μ₁ − μ₂ and standard deviation √(σ₁²/n₁ + σ₂²/n₂). Both are on the formula sheet.
The mean result says x̄₁ − x̄₂ is an unbiased estimator of μ₁ − μ₂. As with proportions, the variances add: each sample brings its own chance variation, so the difference is more variable than either mean alone.
Known σ vs. unknown σ
This topic uses the population standard deviations σ₁ and σ₂, which you'd know only in a textbook setting or from long-run records. With them, x̄₁ − x̄₂ standardizes to a z-score and you can use the normal table. Once you estimate them with s₁ and s₂, you switch to t procedures.
Conditions
- Random: two independent random samples, or random assignment of treatments in an experiment.
- 10%: when sampling without replacement, n₁ ≤ 10% of N₁ and n₂ ≤ 10% of N₂. Not needed for an experiment.
- Shape: if both populations are normal, x̄₁ − x̄₂ is normal. If not, it's approximately normal when n₁ ≥ 30 and n₂ ≥ 30.
Independent groups only
This formula assumes the two samples are independent: who is in one group has nothing to do with who is in the other. Paired data break that assumption. For paired data, work with the differences instead (topic 4.2).
Pairing changes the math because paired values move together. A student who scores high before tends to score high after, so the differences vary much less than the two sets of scores do. The independent-samples formula would overstate the variability.
Interpreting in context
"In repeated random samples of 40 students from School 1 and 50 from School 2, the difference in sample mean scores (School 1 − School 2) would typically vary by about 2.3 points from the true difference of 5 points."
Real problems don't give σ₁ and σ₂, so intervals and tests use s₁ and s₂ and t-distributions (topics 4.7 to 4.10).
What changes the spread
Bigger samples in either group shrink √(σ₁²/n₁ + σ₂²/n₂), but improving the smaller or more variable group helps most. In the example below, School 2 contributes 2.88 to the variance and School 1 only 2.5, because School 2's scores are more spread out.
Compare this with 3.9: the structure is identical. Only the pieces under the square root change, from p(1 − p)/n for proportions to σ²/n for means.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Probability about a difference in means
Scores on a state test have mean 70 and standard deviation 10 at School 1, and mean 65 and standard deviation 12 at School 2. Independent random samples of 40 students from School 1 and 50 from School 2 are selected (each school has over 1,000 students). Find the probability that the School 2 sample mean is higher than the School 1 sample mean.
Show the solutionHide the solution
- Step 1: Mean of x̄₁ − x̄₂: 70 − 65 = 5.
- Step 2: SD: √(10²/40 + 12²/50) = √(2.5 + 2.88) = √5.38 ≈ 2.32.
- Step 3: Shape: both samples have n ≥ 30, so approximately normal. 10%: 40 and 50 are each under 10% of their schools.
- Step 4: "School 2 higher" means x̄₁ − x̄₂ < 0. z = (0 − 5)/2.32 ≈ −2.16.
- Step 5: P(z < −2.16) ≈ 0.0155.
Answer: About 0.016.
- Example 2Calculator allowed
Trap: adding standard deviations
A student computes the standard deviation of x̄₁ − x̄₂ as 10/√40 + 12/√50 ≈ 1.58 + 1.70 = 3.28. What's the correct value?
Show the solutionHide the solution
- Step 1: Standard deviations don't add; variances do.
- Step 2: Var = 10²/40 + 12²/50 = 2.5 + 2.88 = 5.38.
- Step 3: SD = √5.38 ≈ 2.32.
Answer: About 2.32. Add the variances, then take the square root.
Common mistakes
- Adding or subtracting standard deviations instead of adding variances.
- Applying the formula to paired data.
- Checking the sample size for only one of the two groups.
- Forgetting to state the order of subtraction.
On the exam
- "Describe the sampling distribution of x̄₁ − x̄₂": give the shape (with justification), the mean and the standard deviation, in context.
- Watch for questions that give σ values (use the normal model) versus sample data with s values (use t procedures).
Connected topics
Videos
Check yourself
4 questions on 4.6 Sampling Distributions for the Difference Between Two Sample Means. Pick an answer to see if you got it, and why.
Apples from Orchard A have mean weight 150 g (σ = 20 g), and apples from Orchard B have mean weight 140 g (σ = 15 g). Independent random samples of 40 apples from A and 50 apples from B are selected. What are the mean and standard deviation of the sampling distribution of x̄_A − x̄_B?
Using the sampling distribution of x̄_A − x̄_B with mean 10 g and standard deviation 3.81 g (approximately normal), what is the probability that the sample of Orchard B apples has a larger mean weight than the sample of Orchard A apples?
Two populations of test scores are both strongly skewed. Independent random samples of sizes 12 and 15 are taken. Which statement about the sampling distribution of x̄₁ − x̄₂ is correct?
Many pairs of independent random samples are drawn from two populations with means μ₁ = 75 and μ₂ = 82. Which describes the center of the distribution of x̄₁ − x̄₂?
0 of 4 answered