AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3/3-9)
Unit 3 · Topic 3.9
3.9 Sampling Distributions for the Difference Between Sample Proportions
To compare two groups, you look at the difference p̂₁ − p̂₂. This topic gives its sampling distribution: center p₁ − p₂, standard deviation from the formula sheet and the conditions for a normal shape. It's the foundation for the two-sample interval and test that follow.
Key terms
- difference in sample proportions (p̂₁ − p̂₂)
- independent samples
- standard deviation of p̂₁ − p̂₂
- normality condition
Center and spread
For two independent random samples, or two groups from a randomized experiment, the sampling distribution of p̂₁ − p̂₂ has mean p₁ − p₂, so the difference in sample proportions is an unbiased estimator of the difference in population proportions.
Its standard deviation is √(p₁(1 − p₁)/n₁ + p₂(1 − p₂)/n₂). The variances add, even though you're subtracting the proportions. Two sources of chance variation, one from each sample, both make the difference less predictable.
What affects the spread
Larger samples in either group shrink the standard deviation, but the smaller group limits you most. Adding 100 people to a group of 50 helps far more than adding 100 to a group of 1,000.
The standard deviation formula uses the true p₁ and p₂. In real problems they're unknown, so intervals and tests replace them with estimates, giving a standard error (topics 3.10 and 3.12).
Conditions
- Random: two independent random samples, or random assignment of treatments in an experiment. Independent samples means the selection in one group has nothing to do with the other.
- 10%: when sampling without replacement, each sample is at most 10% of its population. Not needed for a randomized experiment.
- Large counts: n₁p₁, n₁(1 − p₁), n₂p₂ and n₂(1 − p₂) are all at least 10. Then p̂₁ − p̂₂ is approximately normal.
Interpreting in context
Name both populations and the order of subtraction. "In repeated random samples, the difference (city minus suburbs) in the proportion of residents who commute by public transit would typically vary by about 0.06 from the true difference."
The order of subtraction is your choice, but stay consistent. Switching it flips the sign of every value.
Independent vs. paired
This formula needs two independent groups. If the same people answer two questions, or the groups are matched pairs, the samples aren't independent and the formula doesn't apply. Watch for this in study descriptions.
Experiments count too
In a randomized experiment, the two groups aren't samples from two populations; they're created by random assignment. The same sampling distribution still works, because random assignment produces chance variation in p̂₁ − p̂₂ much like random sampling does. That's why the random condition can be met either way, and why the 10% condition doesn't apply to experiments: there's no population being sampled without replacement.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Probability of a reversed difference
In Town 1, 60% of adults support a new park; in Town 2, 45% do. Independent random samples of 120 adults from Town 1 and 150 from Town 2 are taken (both towns have over 10,000 adults). Find the probability that the sample shows a smaller proportion supporting the park in Town 1 than in Town 2.
Show the solutionHide the solution
- Step 1: Mean of p̂₁ − p̂₂: 0.60 − 0.45 = 0.15.
- Step 2: SD: √(0.60 × 0.40/120 + 0.45 × 0.55/150) = √(0.0020 + 0.00165) ≈ 0.0604.
- Step 3: Large counts: 120(0.60) = 72, 120(0.40) = 48, 150(0.45) = 67.5, 150(0.55) = 82.5, all ≥ 10, so approximately normal. 10%: each sample is far less than 10% of its town.
- Step 4: "Town 1 smaller" means p̂₁ − p̂₂ < 0. z = (0 − 0.15)/0.0604 ≈ −2.48.
- Step 5: P(z < −2.48) ≈ 0.0066 (technology: 0.0065).
Answer: About 0.0065. A sample in which Town 1 shows less support would be very unusual.
- Example 2Calculator allowed
Trap: subtracting the standard deviations
A student finds σp̂₁ = 0.045 and σp̂₂ = 0.041 and says the standard deviation of p̂₁ − p̂₂ is 0.045 − 0.041 = 0.004. What's the correct value?
Show the solutionHide the solution
- Step 1: Variances add for a difference of independent statistics: σ² = 0.045² + 0.041² = 0.002025 + 0.001681 = 0.003706.
- Step 2: σ = √0.003706 ≈ 0.061.
Answer: About 0.061. Standard deviations don't subtract; the variances add.
Common mistakes
- Subtracting standard deviations or variances instead of adding variances.
- Applying the formula to paired or dependent samples.
- Checking large counts for only one of the two samples.
- Forgetting to state which group is subtracted from which.
On the exam
- Define the order of subtraction once, clearly ("Town 1 minus Town 2"), and use it throughout.
- For an experiment, the random condition is random assignment, and the 10% condition doesn't apply.
Connected topics
- Unit 33.2 Sampling Distributions for Sample Proportions
- Unit 33.10 Constructing a Confidence Interval for the Difference Between Two Population Proportions
- Unit 33.12 Setting Up a Test for the Difference Between Two Population Proportions
- Unit 44.6 Sampling Distributions for the Difference Between Two Sample Means
Videos
Check yourself
3 questions on 3.9 Sampling Distributions for the Difference Between Sample Proportions. Pick an answer to see if you got it, and why.
In a state, 55% of rural adults and 40% of urban adults own a pickup truck. Independent random samples of 200 rural adults and 250 urban adults are selected. What are the mean and standard deviation of the sampling distribution of p̂_rural − p̂_urban?
Which situation meets the conditions for a normal sampling distribution of p̂₁ − p̂₂?
Many pairs of independent random samples are taken from two populations with p₁ = 0.62 and p₂ = 0.48, and p̂₁ − p̂₂ is computed for each pair. Which statement about the distribution of these differences is true?
0 of 3 answered