AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/4/4-4)
Unit 4 · Topic 4.4
4.4 Setting Up a Test for a Population Mean or Population Mean Difference
A one-sample t-test checks a claim about a population mean when σ is unknown. For matched pairs, you test the mean difference μd. Setting up means naming the test, defining the parameter, writing hypotheses and checking the three conditions.
Key terms
- one-sample t-test
- H₀: μ = μ₀
- paired t-test (μd)
- sample data condition
Choosing the test
Use a one-sample t-test for μ when you have one sample of a quantitative variable and want to test a claim about its population mean. Use a paired t-test (a one-sample t-test on the differences) when each unit gives two related measurements or units are matched in pairs.
For inference about means, use t procedures, because σ isn't known. z procedures are for proportions (and for the textbook-style sampling distribution problems in 4.1 and 4.6, where σ is given).
Hypotheses
One mean: H₀: μ = μ₀, with Hₐ: μ < μ₀, μ > μ₀ or μ ≠ μ₀.
Paired: H₀: μd = 0 (no mean difference), with Hₐ: μd < 0, μd > 0 or μd ≠ 0.
Define the parameter in context: "μ = the true mean fill volume of all bottles filled by the machine today" or "μd = the true mean difference (after − before) in practice test score for all students at the school." The order of subtraction decides the direction of Hₐ.
Conditions
For paired data, run the checks on the list of differences.
- Random: a random sample or a randomized experiment (for paired experiments, the order of treatments is often randomized).
- 10%: when sampling without replacement, n ≤ 10% of N.
- Normal/sample data: n ≥ 30; or the population of values (or differences) is known to be roughly normal; or a graph of the sample values (or differences) shows no strong skew or outliers.
Paired or not?
Ask: is each value in one group linked to a specific value in the other? Same subject measured twice, twins, left hand vs. right hand, the same car with two fuel types: paired. Two separate groups of different people: independent (topic 4.9).
Pairing usually helps. Differences within a pair remove person-to-person variation, so a paired test can detect a smaller effect than a two-sample test on the same number of measurements.
Why the sample data condition matters
t procedures hold up well when the population is only roughly normal, as long as there are no outliers or strong skew. With a small sample, though, one outlier can drag x̄ and inflate s, which changes t a lot. That's why you look at a graph of the data when n < 30.
If the graph shows strong skew or an outlier in a small sample, say the condition isn't met and that the results may not be reliable. Don't just quietly continue.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Set up a one-sample t-test
A machine is supposed to fill bottles with 500 mL. A quality inspector suspects underfilling and measures a random sample of 25 bottles from today's run of 8,000. A dotplot of the volumes is roughly symmetric with no outliers. Set up the test.
Show the solutionHide the solution
- Step 1: Test: one-sample t-test for μ.
- Step 2: μ = the true mean volume of all bottles filled by the machine today.
- Step 3: H₀: μ = 500 mL. Hₐ: μ < 500 mL.
- Step 4: Random: random sample of 25 bottles. 10%: 25 ≤ 800 (10% of 8,000). Normal: n < 30, but the dotplot is roughly symmetric with no outliers.
Answer: One-sample t-test with H₀: μ = 500 and Hₐ: μ < 500, where μ is today's true mean fill volume. Conditions are met.
- Example 2Calculator allowed
Set up a paired test
Using the before/after practice test data for 8 randomly selected students from a large school (table in 4.2), set up a test of whether the review session raises scores on average.
Show the solutionHide the solution
- Step 1: Same students measured twice, so the data are paired. Test: paired t-test (one-sample t-test on the differences).
- Step 2: μd = the true mean difference (after − before) in score for all students at the school.
- Step 3: H₀: μd = 0. Hₐ: μd > 0.
- Step 4: Random: randomly selected students. 10%: 8 is less than 10% of a large school. Normal: the 8 differences (12, 5, −3, 8, 10, 6, 0, 9) show no strong skew or outliers.
Answer: Paired t-test with H₀: μd = 0 vs. Hₐ: μd > 0, where μd is the true mean improvement (after − before). Conditions are met.
- Example 3Calculator allowed
Trap: hypotheses about x̄
A student writes H₀: x̄ = 500, Hₐ: x̄ < 497.8 for the bottle test. Fix it.
Show the solutionHide the solution
- Step 1: Hypotheses are about the population mean μ, not the sample mean.
- Step 2: The sample result (497.8) never belongs in the hypotheses.
Answer: H₀: μ = 500 and Hₐ: μ < 500.
Common mistakes
- Using a two-sample test for paired data, or a paired test for independent groups.
- Writing hypotheses with x̄ or with the sample value.
- Checking normality on the before and after values separately instead of on the differences.
- Not stating the order of subtraction for μd.
On the exam
- "Identify the appropriate test" questions often hinge on paired vs. two-sample. Look at how the data were collected.
- Show the sample-data check: "n = 25 < 30, but the dotplot shows no strong skew or outliers."
Connected topics
- Unit 33.5 Setting Up a Test for a Population Proportion
- Unit 44.2 Constructing a Confidence Interval for a Population Mean or Population Mean Difference
- Unit 44.5 Carrying Out a Test for a Population Mean or Population Mean Difference
- Unit 44.9 Setting Up a Test for the Difference Between Two Population Means
Videos
Check yourself
4 questions on 4.4 Setting Up a Test for a Population Mean or Population Mean Difference. Pick an answer to see if you got it, and why.
A random sample of 12 home prices from a town is used to build a t-interval for the mean home price. A dotplot of the 12 prices shows 11 prices between $180,000 and $260,000 and one price of $910,000. Why isn't a t-interval appropriate?
A random sample of 60 delivery times is strongly skewed right. A researcher wants a t-interval for the mean delivery time. Which statement is correct?
A researcher takes a random sample of 40 students, without replacement, from a school of 300 students to estimate the mean time spent on homework. Which condition is NOT met?
A nutritionist wants to test whether the mean sugar content of a brand's breakfast bars is more than the 12 grams printed on the label. Which hypotheses are correct?
0 of 4 answered