Skip to main content

Unit 3 · Topic 3.13

3.13 Carrying Out a Test for the Difference Between Two Population Proportions

To carry out the two-sample z-test, compute z with the pooled standard error, find the p-value from the standard normal distribution and write a conclusion in context about the difference between the populations or treatments.

Key terms

  • z test statistic
  • p-value
  • significance level (α)
  • conclusion in context

Test statistic

z = (p̂₁ − p̂₂ − 0) ÷ √(p̂c(1 − p̂c)(1/n₁ + 1/n₂)).

The 0 is the hypothesized difference. The standard error uses p̂c because the test assumes p₁ = p₂. The formula sheet gives this as the standard error "when p₁ = p₂ is assumed."

p-value and decision

Find the p-value from the standard normal distribution in the direction of Hₐ, using Table A or technology (2-PropZTest). Compare it with α.

For Hₐ: p₁ − p₂ > 0, use the area to the right of z. For Hₐ: p₁ − p₂ < 0, the area to the left. For a two-sided alternative, double the smaller tail area.

Interpret the p-value with the null assumption in context: "Assuming text reminders have no effect on vaccination rates, there's about a 0.013 probability of seeing a difference in sample proportions of 0.10 or more by chance alone."

Conclusion

Comparison, decision, context. "Because the p-value of 0.013 is less than α = 0.05, we reject H₀. There is convincing evidence that the true proportion of patients like these who would get a flu shot is higher with a text reminder than without."

In an experiment, add the cause-and-effect claim if H₀ is rejected. With two random samples, generalize to the populations, but don't claim causation.

If you fail to reject, say there isn't convincing evidence that the proportions differ (in the direction of Hₐ). Never conclude the two proportions are equal.

Connecting to the interval

The 95% interval for the same data was (0.013, 0.187), entirely above 0, and the one-sided test gives p ≈ 0.013. They agree. Note that the test and interval use slightly different standard errors (pooled vs. unpooled), so in borderline cases they can disagree a little.

Possible errors here

If you reject H₀, the possible mistake is a Type I error: concluding reminders help when they actually don't. If you fail to reject, the possible mistake is a Type II error: missing a real benefit. For the flu-shot study, a Type II error might mean a useful, cheap reminder program never gets adopted.

Statistical significance also isn't the whole story. A difference of 10 percentage points in vaccination is large enough to matter for public health. In a huge study, a difference of 1 point could be significant but too small to be worth acting on.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Carry out a two-sample z-test

    Continue the flu-shot test from 3.12: p̂₁ = 0.58, p̂₂ = 0.48, n₁ = n₂ = 250, p̂c = 0.53. Test H₀: p₁ − p₂ = 0 vs. Hₐ: p₁ − p₂ > 0 at α = 0.05.

    Show the solution
    1. Step 1: SE = √(0.53 × 0.47 × (1/250 + 1/250)) ≈ 0.0446.
    2. Step 2: z = (0.58 − 0.48 − 0)/0.0446 ≈ 2.24.
    3. Step 3: p-value = P(z ≥ 2.24) ≈ 0.0125.
    4. Step 4: 0.0125 < 0.05, so reject H₀.

    Answer: z ≈ 2.24, p-value ≈ 0.013. There is convincing evidence that text reminders increase the proportion of patients like these who get a flu shot. Because treatments were randomly assigned, the reminders caused the increase.

  2. Example 2Calculator allowed

    Trap: two-sided p-value

    If the researchers had used Hₐ: p₁ − p₂ ≠ 0 with the same data, what would the p-value be, and would the decision at α = 0.01 change?

    Show the solution
    1. Step 1: Two-sided: double the one tail. p-value ≈ 2 × 0.0125 = 0.025.
    2. Step 2: At α = 0.05: 0.025 < 0.05, still reject.
    3. Step 3: At α = 0.01: 0.025 > 0.01, fail to reject. The choice of α and of a one- or two-sided test can change the decision, which is why both must be chosen before seeing the data.

    Answer: p-value ≈ 0.025. At α = 0.01 you would fail to reject H₀.

Common mistakes

  • Using the unpooled standard error in the test statistic.
  • Doubling the p-value for a one-sided test, or not doubling it for a two-sided test.
  • Claiming causation from an observational comparison of two samples.
  • Ending with "reject H₀" and no statement in context.

On the exam

  • Free-response inference answers are scored on hypotheses, conditions, mechanics and conclusion. Make the conclusion match Hₐ in direction and context.
  • Show the z formula with numbers substituted before the calculator result.

Connected topics

Videos

  • Two-Sample Z Test for Difference in Proportions | AP Statistics (Step-by-Step Guide)

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • Hypothesis test for difference in proportions example | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • AP STATS FRQ 2019 #4 Walkthrough 2 Proportion Z-test

    The AlgebrosWatch on YouTube (opens in a new tab)

  • Inference for Two Proportions: An Example of a Confidence Interval and a Hypothesis Test

    jbstatisticsWatch on YouTube (opens in a new tab)

  • Hypothesis Testing With Two Proportions

    The Organic Chemistry TutorWatch on YouTube (opens in a new tab)

Check yourself

2 questions on 3.13 Carrying Out a Test for the Difference Between Two Population Proportions. Pick an answer to see if you got it, and why.

Question 1 of 2Calculator allowed

A test of H₀: p₁ = p₂ versus Hₐ: p₁ ≠ p₂, where p₁ and p₂ are the proportions of left-handed people in two countries, gives z = 1.21 and a p-value of 0.226. At α = 0.05, which is the correct conclusion?

A clinic randomly assigned 490 patients with upcoming appointments to two groups. The 250 patients in one group got a text reminder the day before, and 185 of them showed up. The 240 patients in the other group got no reminder, and 156 of them showed up.

Described experiment with invented results

Question 2 of 2Calculator allowed

What are the test statistic and p-value for Hₐ: p_text > p_none?

0 of 2 answered