AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3/3-10)
Unit 3 · Topic 3.10
3.10 Constructing a Confidence Interval for the Difference Between Two Population Proportions
A two-sample z-interval estimates the difference between two population proportions, p₁ − p₂. You'll name the procedure and parameter, check conditions for both groups and compute (p̂₁ − p̂₂) ± z*·SE.
Key terms
- two-sample z-interval
- point estimate p̂₁ − p̂₂
- standard error
- margin of error
The formula
(p̂₁ − p̂₂) ± z* √(p̂₁(1 − p̂₁)/n₁ + p̂₂(1 − p̂₂)/n₂).
The point estimate is p̂₁ − p̂₂. The standard error uses the sample proportions, because p₁ and p₂ are unknown. Notice the two pieces under the square root are added: each sample brings its own chance variation, so a difference is less precise than either proportion alone. That's also why a two-sample interval is usually wider than a one-sample interval built from groups of the same size. The margin of error is z* times the standard error. z* values are the same as for one proportion (1.645, 1.960, 2.576 for 90%, 95%, 99%).
The parameter
Define it with both groups and the order: "p₁ − p₂, where p₁ is the true proportion of patients who would get a flu shot if sent a text reminder and p₂ is the true proportion who would get one without a reminder."
For an experiment, the parameters describe what would happen to subjects like these under each treatment. For two random samples, they describe the two populations.
Say which group is subtracted from which before computing anything. A clear definition up front makes the sign of every later number easy to interpret.
Conditions
- Random: two independent random samples or a randomized experiment.
- 10%: when sampling without replacement, n₁ ≤ 10% of N₁ and n₂ ≤ 10% of N₂. Skip it for a randomized experiment.
- Normal: observed successes and failures in each group, n₁p̂₁, n₁(1 − p̂₁), n₂p̂₂ and n₂(1 − p̂₂), are all at least 10.
Reading the result
Because it's a difference, the interval can include negative values, positive values or both. Positive values mean group 1's proportion is higher; negative values mean it's lower. Topic 3.11 covers interpreting and using it.
A complete answer
On a calculator, 2-PropZInt takes the success counts and sample sizes for each group. Enter the groups in the same order as your parameter definition, or the signs of the interval will flip.
If you switch the order of subtraction, the interval becomes (−b, −a). Both versions say the same thing; just describe it correctly.
- Name the procedure: two-sample z-interval for p₁ − p₂.
- Define p₁ and p₂ in context and say which is subtracted.
- Check random, 10% (if sampling) and normal for both groups, with numbers.
- Show the formula with values, give the interval and interpret it in context.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Interval from a randomized experiment
A clinic randomly assigns 500 patients to two groups of 250. One group gets a text reminder to get a flu shot; the other doesn't. 145 of the reminder group and 120 of the no-reminder group get the shot. Construct a 95% confidence interval for the difference in proportions (reminder minus no reminder).
Show the solutionHide the solution
- Step 1: Procedure: two-sample z-interval for p₁ − p₂ (p₁: reminder, p₂: no reminder).
- Step 2: Random: treatments were randomly assigned. 10%: not needed (experiment). Normal: 145, 105, 120 and 130 are all at least 10.
- Step 3: p̂₁ = 145/250 = 0.58, p̂₂ = 120/250 = 0.48. Point estimate: 0.10.
- Step 4: SE = √(0.58 × 0.42/250 + 0.48 × 0.52/250) ≈ 0.0444.
- Step 5: Margin of error = 1.96 × 0.0444 ≈ 0.087.
- Step 6: Interval: 0.10 ± 0.087 = (0.013, 0.187).
Answer: (0.013, 0.187). We are 95% confident that the interval from 0.013 to 0.187 captures the true difference in the proportion of patients like these who would get a flu shot with a text reminder vs. without one.
- Example 2Calculator allowed
Trap: pooling in an interval
A student computes the standard error for the interval above using the combined proportion p̂c = 265/500 = 0.53. Is that right?
Show the solutionHide the solution
- Step 1: Pooling assumes p₁ = p₂, which is the null hypothesis of a test.
- Step 2: An interval doesn't assume the proportions are equal; it estimates how different they are.
- Step 3: So the interval uses each group's own p̂ in the standard error.
Answer: No. Pooling is only for the two-sample z-test. The interval uses p̂₁ and p̂₂ separately.
Common mistakes
- Using the pooled proportion in a confidence interval.
- Checking the normal condition for only one group.
- Requiring the 10% condition for a randomized experiment.
- Defining the parameter as the difference in sample proportions.
On the exam
- Name the procedure exactly ("two-sample z-interval for p₁ − p₂") and define both proportions and the order of subtraction.
- Free response often asks you to check conditions for an experiment. The random condition is random assignment.
Connected topics
- Unit 33.3 Constructing a Confidence Interval for a Population Proportion
- Unit 33.9 Sampling Distributions for the Difference Between Sample Proportions
- Unit 33.11 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions
- Unit 44.7 Constructing a Confidence Interval for the Difference Between Two Population Means
Videos
Check yourself
2 questions on 3.10 Constructing a Confidence Interval for the Difference Between Two Population Proportions. Pick an answer to see if you got it, and why.
Which data collection method allows a two-sample z-interval for p₁ − p₂ to be used, assuming the counts are large enough?
A clinic randomly assigned 490 patients with upcoming appointments to two groups. The 250 patients in one group got a text reminder the day before, and 185 of them showed up. The 240 patients in the other group got no reminder, and 156 of them showed up.
Described experiment with invented results
Which is a 95% confidence interval for p_text − p_none, the difference in show-up rates?
0 of 2 answered