AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3/3-2)
Unit 3 · Topic 3.2
3.2 Sampling Distributions for Sample Proportions
This topic pins down the sampling distribution of p̂: its center, spread and shape, and the conditions that make those facts true. Every proportion interval and test in this unit is built on it.
Key terms
- sampling distribution of p̂
- standard deviation √(p(1 − p)/n)
- randomization condition
- 10% condition
- normality (large counts) condition
Center and spread
If you take a random sample of size n from a population with proportion p, the sampling distribution of p̂ has mean μp̂ = p and standard deviation σp̂ = √(p(1 − p)/n). Both are on the formula sheet.
The mean equaling p is another way of saying p̂ is unbiased. The n in the denominator means larger samples give less variable p̂ values. Because n is under a square root, you need 4 times the sample size to cut the standard deviation in half.
Where the formulas come from
p̂ is just a binomial count divided by n. If X is the number of successes, X is binomial with mean np and standard deviation √(np(1 − p)) (topic 2.10). Dividing by n gives p̂ = X/n, with mean np/n = p and standard deviation √(np(1 − p))/n = √(p(1 − p)/n).
The spread is largest when p = 0.5 and shrinks as p moves toward 0 or 1. A proportion near 0.5 is the hardest to pin down; that's why p̂ = 0.5 is the safe choice when planning a sample size (topic 3.3).
The conditions
Check three conditions before you use these facts. The first two make the mean and standard deviation formulas work; the third makes the shape approximately normal. If the large counts condition fails, the sampling distribution is noticeably skewed and normal calculations will be off.
- Random: the data come from a random sample. This is what makes the mean and standard deviation formulas apply.
- 10% condition: when sampling without replacement, n ≤ 0.10N. This keeps the observations close enough to independent for the standard deviation formula to be accurate.
- Large counts: np ≥ 10 and n(1 − p) ≥ 10. These are the expected numbers of successes and failures. When both are at least 10, the sampling distribution of p̂ is approximately normal.
Using the normal model
When the conditions hold, you can find probabilities about p̂ just as in 2.11, with z = (p̂ − p)/σp̂. This answers questions like "How likely is a sample proportion this far from the truth?" which is the core idea behind significance tests.
Interpret the standard deviation in context: "In repeated random samples of 200 voters, the sample proportion who support the measure typically varies by about 0.032 from the true proportion of 0.30."
Three distributions, again
The population distribution here is just successes and failures. A sample's data is one set of yes/no answers. The sampling distribution of p̂ is the distribution of the proportion across all possible samples. Only the sampling distribution has the mean p and standard deviation √(p(1 − p)/n).
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Probability about p̂
30% of a city's registered voters support a ballot measure. A pollster takes a random sample of 200 of the city's 85,000 voters. Describe the sampling distribution of p̂ and find the probability that more than 35% of the sample supports the measure.
Show the solutionHide the solution
- Step 1: Mean: μp̂ = p = 0.30.
- Step 2: Standard deviation: √(0.30 × 0.70/200) = √0.00105 ≈ 0.0324. This is valid since 200 ≤ 10% of 85,000 = 8,500.
- Step 3: Shape: np = 200(0.30) = 60 and n(1 − p) = 140, both at least 10, so approximately normal.
- Step 4: z = (0.35 − 0.30)/0.0324 ≈ 1.54. P(z > 1.54) ≈ 0.061 (technology: 0.0614).
Answer: p̂ is approximately normal with mean 0.30 and standard deviation about 0.032. P(p̂ > 0.35) ≈ 0.061.
- Example 2Calculator allowed
Trap: the large counts condition fails
With the same p = 0.30, a student takes a random sample of n = 20 and uses a normal model to find P(p̂ > 0.50). Is that valid?
Show the solutionHide the solution
- Step 1: Check large counts: np = 20(0.30) = 6, which is less than 10.
- Step 2: The sampling distribution of p̂ for n = 20 is noticeably skewed right, so normal probabilities aren't reliable.
- Step 3: A binomial calculation (2.10) on the count X = 20p̂ would be the right tool here.
Answer: No. np = 6 < 10, so the sampling distribution of p̂ isn't approximately normal for n = 20.
Common mistakes
- Using p̂ in the standard deviation formula when p is known. σp̂ uses the true p.
- Checking large counts with the sample size alone ("n is bigger than 30"). For proportions, check np and n(1 − p).
- Saying the 10% condition makes the distribution normal. It's about independence; large counts is about shape.
- Saying "the sample is approximately normal" instead of "the sampling distribution of p̂ is approximately normal."
On the exam
- "Describe the sampling distribution" means shape, center and spread, each with a justification: normal because of large counts, mean p, standard deviation from the formula with the 10% condition.
- Show the numbers when checking conditions: "np = 60 ≥ 10 and n(1 − p) = 140 ≥ 10."
Connected topics
Videos
Check yourself
4 questions on 3.2 Sampling Distributions for Sample Proportions. Pick an answer to see if you got it, and why.
In a large city, 40% of households take part in a curbside composting program. A random sample of 150 households is selected, and p̂ is the proportion of sampled households that take part.
Described scenario with an invented percent
What are the mean and standard deviation of the sampling distribution of p̂?
Is the sampling distribution of p̂ approximately normal?
What is the approximate probability that more than 45% of the sampled households take part in the program?
For random samples of size n = 100, the standard deviation of the sampling distribution of p̂ is 0.05. What is the standard deviation for random samples of size n = 400 from the same population?
0 of 4 answered