AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/4/4-1)
Unit 4 · Topic 4.1
4.1 Sampling Distributions for Sample Means
The sample mean x̄ varies from sample to sample, and its sampling distribution has a predictable center, spread and shape. Its mean is μ, its standard deviation is σ/√n, and it's normal or approximately normal under the right conditions. This is the basis for every t procedure in this unit.
Key terms
- sampling distribution of x̄
- mean of x̄ equals μ
- standard deviation σ/√n
- central limit theorem
- 10% condition
Center and spread
For a random sample of size n from a population with mean μ and standard deviation σ, the sampling distribution of x̄ has mean μx̄ = μ and standard deviation σx̄ = σ/√n. Both are on the formula sheet.
μx̄ = μ means x̄ is an unbiased estimator of μ. The √n in the denominator means averages are less variable than individual values, and they get less variable as n grows. To cut the standard deviation in half, you need four times the sample size.
How sample size shrinks the spread
With σ = 12 minutes: samples of 9 give σx̄ = 12/3 = 4 minutes, samples of 36 give 12/6 = 2 minutes, and samples of 144 give 12/12 = 1 minute. Each time n is multiplied by 4, the spread of x̄ is cut in half.
The spread of the individual values doesn't change with n. A bigger sample doesn't make people's wait times less variable; it makes the average of those wait times more predictable.
In practice σ is rarely known. When you replace it with s, you'll use t-distributions instead of the normal (topic 4.2).
Conditions for the formulas
- Random: the data come from a random sample.
- 10%: when sampling without replacement, n ≤ 10% of N. This keeps the observations close enough to independent for σ/√n to be accurate.
Shape
If the population is normal, the sampling distribution of x̄ is exactly normal for any sample size, even n = 2.
If the population isn't normal, the central limit theorem says x̄ is approximately normal when n is large; the usual guideline is n ≥ 30. A strongly skewed population (like incomes) may need a sample much larger than 30.
If the population isn't normal and n is small, the sampling distribution of x̄ will share some of the population's skew, so normal calculations aren't safe.
Individuals vs. means
Read questions carefully. "The probability that one randomly chosen customer waits more than 43 minutes" uses the population distribution, with σ. "The probability that the mean wait for 36 customers is more than 43 minutes" uses the sampling distribution, with σ/√n.
Means are much less spread out than individuals, so an extreme sample mean is much rarer than an extreme individual value.
Interpreting in context
"In repeated random samples of 36 customers, the sample mean wait time typically varies by about 2 minutes from the true mean of 40 minutes." Always name the statistic, the sample size and the population.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Probability about a mean (CLT)
Wait times at a busy clinic are skewed right with mean 40 minutes and standard deviation 12 minutes. A random sample of 36 patients is selected from the thousands seen each month. Find the probability that their mean wait is more than 43 minutes.
Show the solutionHide the solution
- Step 1: Mean of x̄: 40 minutes.
- Step 2: Standard deviation of x̄: 12/√36 = 2 minutes (36 is less than 10% of the patients).
- Step 3: Shape: the population is skewed, but n = 36 ≥ 30, so by the central limit theorem x̄ is approximately normal.
- Step 4: z = (43 − 40)/2 = 1.5. P(z > 1.5) ≈ 0.0668.
Answer: About 0.067.
- Example 2Calculator allowed
Normal population, small sample
Bags of chips have weights that are normally distributed with mean 500 g and standard deviation 8 g. An inspector weighs a random sample of 4 bags. Find the probability that their mean weight is less than 495 g.
Show the solutionHide the solution
- Step 1: The population is normal, so x̄ is normal even with n = 4.
- Step 2: σx̄ = 8/√4 = 4 g.
- Step 3: z = (495 − 500)/4 = −1.25. P(z < −1.25) ≈ 0.1056.
Answer: About 0.106.
- Example 3Calculator allowed
Trap: one individual from a skewed population
Using the clinic data (skewed right, μ = 40, σ = 12), a student computes the probability that a single patient waits more than 43 minutes as P(z > 0.25) ≈ 0.40. What's wrong?
Show the solutionHide the solution
- Step 1: For one patient, you need the population distribution, which is skewed right, not normal.
- Step 2: The central limit theorem applies to sample means, not to individual values.
- Step 3: Without knowing the population's exact shape, you can't find this probability with a normal model.
Answer: The normal model doesn't apply to one patient's wait, because the population is skewed. The CLT only helps for the mean of a large sample.
Common mistakes
- Using σ instead of σ/√n for a probability about a sample mean.
- Applying the central limit theorem to individual values or to the sample's data.
- Assuming n ≥ 30 is always enough, even for very strongly skewed populations.
- Saying "the sample is approximately normal" instead of "the sampling distribution of x̄ is approximately normal."
On the exam
- When you use the normal model for x̄, justify the shape: "the population is normal" or "n = 36 ≥ 30, so by the central limit theorem…"
- Questions often pair an individual probability with a sample-mean probability. Decide which distribution each one needs.
Connected topics
Videos
Check yourself
4 questions on 4.1 Sampling Distributions for Sample Means. Pick an answer to see if you got it, and why.
The lifetimes of a brand of rechargeable battery have mean μ = 40 hours and standard deviation σ = 6 hours. A random sample of 36 batteries is selected, and x̄ is the sample mean lifetime.
Described scenario with invented values
What are the mean and standard deviation of the sampling distribution of x̄?
What is the approximate probability that the mean lifetime of the 36 batteries is less than 38 hours?
Why is it reasonable to use a normal model for x̄ here, even though the shape of the population of battery lifetimes is unknown?
The amounts customers spend at a store are strongly skewed right. A manager takes random samples of 8 receipts and computes x̄ for each. Which describes the sampling distribution of x̄?
0 of 4 answered