AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3/3-13)
Unit 3 · Topic 3.13
3.13 Carrying Out a Test for the Difference Between Two Population Proportions
To carry out the two-sample z-test, compute z with the pooled standard error, find the p-value from the standard normal distribution and write a conclusion in context about the difference between the populations or treatments.
Key terms
- z test statistic
- p-value
- significance level (α)
- conclusion in context
Test statistic
z = (p̂₁ − p̂₂ − 0) ÷ √(p̂c(1 − p̂c)(1/n₁ + 1/n₂)).
The 0 is the hypothesized difference. The standard error uses p̂c because the test assumes p₁ = p₂. The formula sheet gives this as the standard error "when p₁ = p₂ is assumed."
p-value and decision
Find the p-value from the standard normal distribution in the direction of Hₐ, using Table A or technology (2-PropZTest). Compare it with α.
For Hₐ: p₁ − p₂ > 0, use the area to the right of z. For Hₐ: p₁ − p₂ < 0, the area to the left. For a two-sided alternative, double the smaller tail area.
Interpret the p-value with the null assumption in context: "Assuming text reminders have no effect on vaccination rates, there's about a 0.013 probability of seeing a difference in sample proportions of 0.10 or more by chance alone."
Conclusion
Comparison, decision, context. "Because the p-value of 0.013 is less than α = 0.05, we reject H₀. There is convincing evidence that the true proportion of patients like these who would get a flu shot is higher with a text reminder than without."
In an experiment, add the cause-and-effect claim if H₀ is rejected. With two random samples, generalize to the populations, but don't claim causation.
If you fail to reject, say there isn't convincing evidence that the proportions differ (in the direction of Hₐ). Never conclude the two proportions are equal.
Connecting to the interval
The 95% interval for the same data was (0.013, 0.187), entirely above 0, and the one-sided test gives p ≈ 0.013. They agree. Note that the test and interval use slightly different standard errors (pooled vs. unpooled), so in borderline cases they can disagree a little.
Possible errors here
If you reject H₀, the possible mistake is a Type I error: concluding reminders help when they actually don't. If you fail to reject, the possible mistake is a Type II error: missing a real benefit. For the flu-shot study, a Type II error might mean a useful, cheap reminder program never gets adopted.
Statistical significance also isn't the whole story. A difference of 10 percentage points in vaccination is large enough to matter for public health. In a huge study, a difference of 1 point could be significant but too small to be worth acting on.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Carry out a two-sample z-test
Continue the flu-shot test from 3.12: p̂₁ = 0.58, p̂₂ = 0.48, n₁ = n₂ = 250, p̂c = 0.53. Test H₀: p₁ − p₂ = 0 vs. Hₐ: p₁ − p₂ > 0 at α = 0.05.
Show the solutionHide the solution
- Step 1: SE = √(0.53 × 0.47 × (1/250 + 1/250)) ≈ 0.0446.
- Step 2: z = (0.58 − 0.48 − 0)/0.0446 ≈ 2.24.
- Step 3: p-value = P(z ≥ 2.24) ≈ 0.0125.
- Step 4: 0.0125 < 0.05, so reject H₀.
Answer: z ≈ 2.24, p-value ≈ 0.013. There is convincing evidence that text reminders increase the proportion of patients like these who get a flu shot. Because treatments were randomly assigned, the reminders caused the increase.
- Example 2Calculator allowed
Trap: two-sided p-value
If the researchers had used Hₐ: p₁ − p₂ ≠ 0 with the same data, what would the p-value be, and would the decision at α = 0.01 change?
Show the solutionHide the solution
- Step 1: Two-sided: double the one tail. p-value ≈ 2 × 0.0125 = 0.025.
- Step 2: At α = 0.05: 0.025 < 0.05, still reject.
- Step 3: At α = 0.01: 0.025 > 0.01, fail to reject. The choice of α and of a one- or two-sided test can change the decision, which is why both must be chosen before seeing the data.
Answer: p-value ≈ 0.025. At α = 0.01 you would fail to reject H₀.
Common mistakes
- Using the unpooled standard error in the test statistic.
- Doubling the p-value for a one-sided test, or not doubling it for a two-sided test.
- Claiming causation from an observational comparison of two samples.
- Ending with "reject H₀" and no statement in context.
On the exam
- Free-response inference answers are scored on hypotheses, conditions, mechanics and conclusion. Make the conclusion match Hₐ in direction and context.
- Show the z formula with numbers substituted before the calculator result.
Connected topics
- Unit 33.7 Carrying Out a Test for a Population Proportion
- Unit 33.11 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions
- Unit 33.12 Setting Up a Test for the Difference Between Two Population Proportions
- Unit 33.8 Potential Errors When Performing Tests
Videos
Check yourself
2 questions on 3.13 Carrying Out a Test for the Difference Between Two Population Proportions. Pick an answer to see if you got it, and why.
A test of H₀: p₁ = p₂ versus Hₐ: p₁ ≠ p₂, where p₁ and p₂ are the proportions of left-handed people in two countries, gives z = 1.21 and a p-value of 0.226. At α = 0.05, which is the correct conclusion?
A clinic randomly assigned 490 patients with upcoming appointments to two groups. The 250 patients in one group got a text reminder the day before, and 185 of them showed up. The 240 patients in the other group got no reminder, and 156 of them showed up.
Described experiment with invented results
What are the test statistic and p-value for Hₐ: p_text > p_none?
0 of 2 answered