Skip to main content

Unit 2 · Topic 2.12

2.12 Sampling Distributions and the Central Limit Theorem

A statistic like x̄ or p̂ changes from sample to sample, and its sampling distribution describes that variation. You'll build sampling distributions and randomization distributions by simulation and meet the central limit theorem, the reason normal models work for sample means. This is the bridge to all of the inference in Units 3 and 4.

Key terms

  • sampling distribution
  • simulation
  • randomization distribution
  • central limit theorem (CLT)

Three different distributions

Keep these apart. Most mistakes in inference start here:

DistributionWhat it showsExample
PopulationValues of a variable for every unit in the populationHeights of all 2,000 students
Sample dataValues for the units in one sampleHeights of the 40 students you measured
Sampling distributionValues of a statistic from every possible sample of size nMean heights from all possible samples of 40 students

Sampling distributions by simulation

You usually can't list every possible sample, so you simulate. Assume a value for the parameter, generate many random samples of size n from that population, compute the statistic for each and graph the results. The dotplot or histogram of those statistics approximates the sampling distribution.

Two patterns show up every time. The sampling distribution is centered near the parameter (for x̄ and p̂). And it gets less spread out as n gets larger: means of samples of 100 vary less than means of samples of 10.

The reason is averaging. In a big sample, unusually high and unusually low values tend to cancel out, so the mean rarely lands far from μ. In a small sample, one or two extreme values can drag the mean a long way. Topics 3.2 and 4.1 give formulas for exactly how fast the spread shrinks.

Randomization distributions

For an experiment, you can ask what results would look like if the treatment had no effect. Write each subject's response on a card, shuffle the cards and deal them back into groups of the original sizes, then compute the statistic, like the difference in group means. Repeat many times.

This randomization distribution shows how big a difference random assignment alone tends to produce. If the actual difference is far out in the tail, chance alone is a poor explanation. That idea becomes the p-value in 3.6.

The central limit theorem

The central limit theorem (CLT) says that for random samples, the sampling distribution of x̄ is approximately normal when n is large enough, even if the population isn't normal. The bigger n is, the closer to normal it gets.

If the population is already normal, x̄ is normal for any n. If the population is skewed, x̄ needs a bigger sample before it looks normal; n ≥ 30 is the usual guide, and strongly skewed populations may need more.

The CLT is about the distribution of x̄, not the data. A large sample from a skewed population is still skewed; it's the means from many samples that pile up in a bell shape.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1Calculator allowed

    Read a simulated sampling distribution

    A school says 40% of its students bike to school. A student simulates 200 random samples of 50 students, assuming p = 0.40, and records p̂ for each. The dotplot is centered at about 0.40 and ranges from about 0.24 to 0.58; only 3 of the 200 values are 0.24 or less. In a real random sample of 50, p̂ = 0.24. What does the simulation suggest?

    Show the solution
    1. Step 1: The simulation shows the values p̂ typically takes when p really is 0.40.
    2. Step 2: Only 3 of 200 simulated samples (1.5%) gave p̂ ≤ 0.24.
    3. Step 3: A result this low would be rare if p = 0.40, so the real sample is evidence that the proportion of bikers is lower than 40%.

    Answer: A p̂ as low as 0.24 happened in only about 1.5% of simulated samples, so the claim of 40% seems unlikely to be correct.

  2. Example 2Calculator allowed

    Trap: what the CLT is about

    Household incomes in a city are strongly skewed right. A student says, "If we take a random sample of 500 households, the incomes in the sample will be approximately normal by the central limit theorem." Correct the statement.

    Show the solution
    1. Step 1: A large random sample tends to look like the population, so the 500 incomes will still be skewed right.
    2. Step 2: The CLT describes the sampling distribution of the sample mean: the means from many samples of 500 households.
    3. Step 3: With n = 500, that distribution of x̄ will be approximately normal.

    Answer: The sample incomes will be skewed right, like the population. It's the sampling distribution of x̄ (for samples of 500) that is approximately normal.

Common mistakes

  • Saying the sample data become normal as n grows. The CLT is about the sampling distribution of x̄.
  • Mixing up the population distribution, the sample's data and the sampling distribution, often by using the word "it."
  • Saying "larger samples have less variability" without saying you mean the variability of the statistic.
  • Thinking a randomization distribution shows the treatment effect. It shows what happens when there is no effect.

On the exam

  • Name the distribution you mean: "the sampling distribution of the sample mean," not "the distribution" or "it." Readers watch for this.
  • Expect questions with a dotplot of simulated statistics: use it to judge whether an observed result is unusual.

Connected topics

Videos

  • AP Statistics Topic 2.12 Sampling Distributions and Central Limit Theorem | Lesson+Notes+Practice

    Michael Porinchak - AP Statistics & AP PrecalculusWatch on YouTube (opens in a new tab)

  • Introduction to sampling distributions | Sampling distributions | AP Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

  • The Central Limit Theorem, Clearly Explained!!!

    StatQuest with Josh StarmerWatch on YouTube (opens in a new tab)

  • Sampling Distributions (7.2)

    Simple Learning ProWatch on YouTube (opens in a new tab)

  • Introduction to the Central Limit Theorem

    jbstatisticsWatch on YouTube (opens in a new tab)

  • Central limit theorem | Inferential statistics | Probability and Statistics | Khan Academy

    Khan AcademyWatch on YouTube (opens in a new tab)

Check yourself

5 questions on 2.12 Sampling Distributions and the Central Limit Theorem. Pick an answer to see if you got it, and why.

Question 1 of 5Calculator allowed

The amount of time people spend at a museum is strongly skewed right, with mean 95 minutes. A researcher takes one random sample of 400 visitors. Which statement best describes the distribution of the 400 times in this sample?

Question 2 of 5Calculator allowed

A population of commute times is strongly skewed right. Many random samples of size 60 are taken, and the mean of each sample is recorded. Which best describes the distribution of these sample means?

Question 3 of 5Calculator allowed

A simulation draws 1,000 random samples of size n = 10 from a population and makes a dotplot of the 1,000 sample means. The simulation is repeated with n = 100. How does the second dotplot compare with the first?

In an experiment, 20 students were randomly assigned to study a list of words with either a new memory method or the usual method, 10 students in each group. The new-method group recalled an average of 4.2 more words than the usual-method group.

To judge whether a difference this large could happen by chance, a teacher reshuffled the 20 scores into two groups of 10 at random 200 times and recorded the difference in means (new − old) each time. The 200 differences were centered at about 0, and only 3 of them were 4.2 or greater.

Described experiment with invented results

Question 4 of 5Calculator allowed

Which statement describes what each dot in the randomization distribution represents?

Question 5 of 5Calculator allowed

Based on the randomization distribution, which conclusion is most reasonable?

0 of 5 answered