Skip to main content

Unit 4 · Topic 4.10

4.10 Carrying Out a Test for the Difference Between Two Population Means

To carry out a two-sample t-test, compute t = (x̄₁ − x̄₂ − 0)/√(s₁²/n₁ + s₂²/n₂), get the p-value from a t-distribution (usually with technology), compare it with α and write a conclusion in context about the two population means.

Key terms

  • t test statistic
  • p-value
  • significance level (α)
  • conclusion in context

The test statistic

t = (x̄₁ − x̄₂ − 0) ÷ √(s₁²/n₁ + s₂²/n₂). The 0 is the hypothesized difference from H₀. The denominator is the same standard error used in the interval; for means, there's no pooling.

When H₀ is true, t has approximately a t-distribution. Technology (2-SampTTest, with Pooled: No) gives the df, which falls between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2. By hand, using the smaller of n₁ − 1 and n₂ − 1 is the safe choice.

p-value and decision

The p-value is the area beyond t in the direction of Hₐ. Interpret it with the null assumption: "Assuming the true mean commute times in the two cities are equal, there is about a 0.016 probability of getting a difference in sample means of 4.3 minutes or more by chance alone."

If the p-value ≤ α, reject H₀: there is convincing evidence for Hₐ. Otherwise, fail to reject: there isn't convincing evidence for Hₐ.

Conclusion and scope

Write the conclusion about the population means, in context, matching Hₐ: "There is convincing evidence that the true mean commute time for all workers in City A is greater than for all workers in City B."

With two random samples, you can generalize to the two populations but can't claim the city causes longer commutes. With a randomized experiment, rejecting H₀ supports cause and effect for units like those in the study.

Errors and power

Rejecting H₀ risks a Type I error (concluding the means differ when they don't). Failing to reject risks a Type II error (missing a real difference). Larger samples, less variable data, a larger true difference or a larger α all increase power, just as in topic 3.8.

Matching the interval

A two-sided test at α = 0.05 and a 95% two-sample t-interval usually agree: the test rejects when 0 is outside the interval. A one-sided test at α = 0.05 matches a 90% interval instead, because it puts all 5% in one tail.

Reading calculator output

2-SampTTest reports t, the p-value, df, both sample means and both standard deviations. On free response, copy the important numbers into a sentence and show the formula; "2-SampTTest gave p = 0.016" alone usually doesn't earn the mechanics point.

Make sure the order you entered the groups matches your hypotheses. Entering City B first would flip the sign of t and the direction of the test.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Carry out a two-sample t-test

    Continue the commute test from 4.9. City A: n = 35, x̄ = 28.4 min, s = 9.2 min. City B: n = 40, x̄ = 24.1 min, s = 7.5 min. Test H₀: μA − μB = 0 vs. Hₐ: μA − μB > 0 at α = 0.05.

    Show the solution
    1. Step 1: SE = √(9.2²/35 + 7.5²/40) ≈ 1.956.
    2. Step 2: t = (28.4 − 24.1 − 0)/1.956 ≈ 2.20.
    3. Step 3: Technology: df ≈ 65.7, p-value ≈ 0.016. (With the conservative df = 34, p ≈ 0.017.)
    4. Step 4: 0.016 < 0.05, so reject H₀.

    Answer: t ≈ 2.20, df ≈ 65.7, p-value ≈ 0.016. There is convincing evidence that the true mean commute time for all workers in City A is greater than for all workers in City B.

  2. Example 2Calculator allowed

    Trap: claiming causation from samples

    After the commute test, a student writes: "Living in City A causes people to have longer commutes." Evaluate this.

    Show the solution
    1. Step 1: The data come from random samples, not an experiment; nobody was assigned to a city.
    2. Step 2: Other variables could explain the difference, such as city size, traffic or job locations.
    3. Step 3: The test supports a difference in population means, not a cause.

    Answer: Not justified. Random sampling lets you generalize the difference to both cities' workers, but without random assignment you can't conclude the city causes longer commutes.

Common mistakes

  • Using a pooled standard error or the pooled calculator option.
  • Using df = n₁ + n₂ − 2 by hand.
  • Writing the conclusion about sample means or about individuals rather than population means.
  • Claiming causation without random assignment.

On the exam

  • A complete answer shows the test statistic formula with values, t, df and the p-value, then a conclusion linked to α in context.
  • Inference free-response questions often add a part about scope or errors. Be ready to say whether cause and effect or generalization is justified.

Connected topics

Videos

  • AP Statistics | Two Sample t-Test for the Difference Between Two Means (Step-by-Step)

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • Two-sample t test for difference of means | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • AP STATS FRQ 2010 #5 Walkthrough 2 Sample T-Test

    The AlgebrosWatch on YouTube (opens in a new tab)

  • Welch (Unpooled Variance) t Tests and Confidence Intervals: An Example

    jbstatisticsWatch on YouTube (opens in a new tab)

  • Conclusion for a two-sample t test using a P-value | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 4.10 Carrying Out a Test for the Difference Between Two Population Means. Pick an answer to see if you got it, and why.

Question 1 of 4Calculator allowed

A randomized experiment compares mean plant growth with a new fertilizer and with a standard fertilizer. A one-sided test of H₀: μ_new = μ_standard versus Hₐ: μ_new > μ_standard gives p-value = 0.008. Which conclusion is correct at α = 0.05?

Question 2 of 4Calculator allowed

Independent random samples of students who do and don't play a musical instrument show a mean GPA that is significantly higher for instrument players (p-value = 0.002). Which conclusion is appropriate?

Question 3 of 4Calculator allowed

A two-sample t-test comparing two means gives a p-value of 0.09. A researcher repeats the study with much larger samples, and the sample means and standard deviations come out about the same as before. What is most likely to happen to the p-value?

Question 4 of 4Calculator allowed

A two-sample t-test of H₀: μ₁ = μ₂ versus Hₐ: μ₁ > μ₂ gives a p-value of 0.03. Which is the correct interpretation?

0 of 4 answered