Skip to main content

Unit 4 · Topic 4.9

4.9 Setting Up a Test for the Difference Between Two Population Means

A two-sample t-test checks whether two population means differ, using data from two independent random samples or a randomized experiment. Setting it up means confirming the groups are independent (not paired), writing hypotheses about μ₁ − μ₂ and checking conditions for both groups.

Key terms

  • two-sample t-test
  • H₀: μ₁ = μ₂
  • independent vs. paired samples
  • conditions for a test

Hypotheses

H₀: μ₁ = μ₂ (equivalently μ₁ − μ₂ = 0). Hₐ: μ₁ > μ₂, μ₁ < μ₂ or μ₁ ≠ μ₂ (equivalently μ₁ − μ₂ > 0, < 0 or ≠ 0).

Define both means in context with the populations or treatments, and say which is group 1. The direction of Hₐ comes from the research question, set before you see the data.

In an experiment, H₀ says the treatments have the same effect on the mean response.

Conditions

  • Random: two independent random samples or a randomized experiment.
  • 10%: when sampling without replacement, each sample is at most 10% of its population. Not needed for an experiment.
  • Normal/sample data: n₁ ≥ 30 and n₂ ≥ 30, or both populations approximately normal. If either sample is under 30, graphs of both samples show no strong skew or outliers.

Why random assignment matters here

In an experiment, random assignment does two jobs. It makes the groups similar on everything except the treatment, so a significant difference can be blamed on the treatment. And it creates the chance variation that the t-distribution models, which is why it satisfies the random condition even when the subjects are volunteers.

Two-sample or paired?

This is the most important decision in the unit. Use a two-sample t-test when there are two separate groups of different units, with no natural link between a value in one group and a value in the other. Use a paired t-test (4.4) when each value in one group is matched to a specific value in the other.

Quick check: could you line up the two columns of data so each row belongs to the same person, the same pair of twins or the same object? If yes, and the matching was part of the design, it's paired. If the groups even have different sizes, it can't be paired.

DesignProcedure
30 volunteers randomly split into two diet groupsTwo-sample t-test
Each volunteer tries both diets in random orderPaired t-test
Random samples of 40 city and 40 rural householdsTwo-sample t-test
Left-hand and right-hand grip strength for 25 peoplePaired t-test

Choosing among all the tests

Categorical response with one group: one-proportion z-test. Two groups: two-proportion z-test. More than two groups or categories: chi-square. Quantitative response with one group: one-sample t-test. Paired data: paired t-test. Two independent groups: two-sample t-test.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Set up a two-sample t-test

    A researcher wants to know whether workers in City A have a longer mean commute than workers in City B. She takes independent random samples of 35 workers from City A and 40 from City B. Set up the test.

    Show the solution
    1. Step 1: Two independent random samples, quantitative response: two-sample t-test for μA − μB.
    2. Step 2: μA = true mean commute time of all workers in City A; μB = the same for City B.
    3. Step 3: H₀: μA − μB = 0. Hₐ: μA − μB > 0.
    4. Step 4: Random: independent random samples. 10%: 35 and 40 are less than 10% of the workers in each city. Normal: both samples have n ≥ 30.

    Answer: Two-sample t-test with H₀: μA − μB = 0 vs. Hₐ: μA − μB > 0. Conditions are met.

  2. Example 2Calculator allowed

    Trap: choosing the wrong test

    To test whether a new running shoe improves 5K times, a coach has each of 12 runners race once in old shoes and once in new shoes, in random order. A student plans a two-sample t-test. What should the coach use, and why?

    Show the solution
    1. Step 1: Each runner produces two times, one per shoe. The data are paired by runner.
    2. Step 2: Runners differ a lot in speed. Pairing removes that runner-to-runner variation.
    3. Step 3: Use a paired t-test on the 12 differences (old − new), with H₀: μd = 0 and Hₐ: μd > 0 if positive differences mean the new shoes are faster.

    Answer: A paired t-test on the 12 differences, because each runner used both shoes.

Common mistakes

  • Using a two-sample test on paired data (or the reverse).
  • Writing hypotheses about sample means.
  • Checking the sample data condition for only one group.
  • Requiring the 10% condition in a randomized experiment.

On the exam

  • Free-response inference questions often reward naming the correct test. Justify paired vs. two-sample from the design in one sentence.
  • State hypotheses in symbols and define both parameters in words, including the order of subtraction.

Connected topics

Videos

  • AP Stats 8.4 - Hypothesis Test for Two Means

    Skew The ScriptWatch on YouTube (opens in a new tab)

  • Hypotheses for a two-sample t test | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • AP Statistics Inference for Means – Significance Tests for a Difference in Means

    Goldie's Math EmporiumWatch on YouTube (opens in a new tab)

  • Example of hypotheses for paired and two-sample t tests | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • Conditions for inference for difference of means | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 4.9 Setting Up a Test for the Difference Between Two Population Means. Pick an answer to see if you got it, and why.

Question 1 of 4Calculator allowed

In which study should a paired t procedure be used?

Question 2 of 4Calculator allowed

A sports scientist wants to know whether a new running shoe lowers the mean time to run 1 mile. She randomly assigns 40 runners to the new shoe and 40 to their usual shoe. Which hypotheses are correct?

Question 3 of 4Calculator allowed

A researcher records the reaction times of a random sample of 50 teenagers and an independent random sample of 50 adults over 60. She wants to know whether the mean reaction times differ. Which procedure should she use?

Question 4 of 4Calculator allowed

A researcher takes a random sample of 25 laptops of Brand X and an independent random sample of 25 laptops of Brand Y from a large store's stock and records battery life. Dotplots of both samples are roughly symmetric with no outliers. Which statement about the conditions for a two-sample t-test is correct?

0 of 4 answered