AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3/3-6)
Unit 3 · Topic 3.6
3.6 p-Values
The p-value is the probability, if H₀ were true, of getting a result at least as extreme as the one you saw. It measures how surprising your data would be under the null hypothesis: the smaller it is, the stronger the evidence for Hₐ. Interpreting it correctly is one of the most tested skills in the course.
Key terms
- p-value
- null distribution
- test statistic
- convincing evidence
The null distribution
If H₀ is true, the test statistic has a known distribution, called the null distribution. For a one-proportion z-test, it's approximately the standard normal distribution. You can also simulate it: generate many samples assuming p = p₀ and record the statistic each time.
Which tail?
"As extreme or more extreme" means in the direction of Hₐ:
- Hₐ: p > p₀: the area at or above the observed z.
- Hₐ: p < p₀: the area at or below the observed z.
- Hₐ: p ≠ p₀: both tails, the area at or below −|z| plus the area at or above |z|. For a symmetric distribution, that's double one tail.
p-values from a simulation
If the null distribution was simulated, the p-value is the proportion of simulated results at least as extreme as the observed one, in the direction of Hₐ. If 7 of 500 simulated samples gave a p̂ as low as yours, the p-value is about 7/500 = 0.014.
Interpreting a p-value
A complete interpretation has three parts: the assumption that H₀ is true (in context), the probability, and the observed result "or more extreme."
Example: "Assuming the true proportion of on-time deliveries is 0.90, there is about a 0.015 probability of getting a sample proportion of 0.847 or lower in a random sample of 150 deliveries by chance alone."
What a p-value is not
It's not the probability that H₀ is true. It's computed assuming H₀ is true, so it can't measure that.
A large p-value doesn't show H₀ is true. It only means the data are consistent with H₀, and they might also be consistent with many other values. Lack of evidence against H₀ is not evidence for it.
Small p-values mean the observed result would be unusual if H₀ were true, which gives evidence for Hₐ. The smaller the p-value, the more convincing that evidence.
Sample size changes the p-value
The same sample proportion can give very different p-values depending on n. Testing H₀: p = 0.5 vs. Hₐ: p > 0.5 with p̂ = 0.56: with n = 200, z ≈ 1.70 and the p-value is about 0.045. With n = 800, z ≈ 3.39 and the p-value is about 0.0003.
Bigger samples make the sampling distribution narrower, so the same distance from p₀ is more surprising. With a very large sample, even a tiny difference can be statistically significant, which is why you should also ask whether the difference matters in practice.
Worked examples
Try each one yourself first, then open the solution.
- Example 1Calculator allowed
Interpret a p-value
A test of H₀: p = 0.50 vs. Hₐ: p > 0.50, where p is the proportion of all students at a school who prefer a later start, gives p̂ = 0.56 from a random sample of 200 students and a p-value of 0.045. Interpret the p-value.
Show the solutionHide the solution
- Step 1: Start with the assumption: the true proportion is 0.50.
- Step 2: State the probability and the result in the direction of Hₐ: 0.56 or higher.
- Step 3: Mention the sample size and chance variation.
Answer: Assuming 50% of all students at the school prefer a later start, there is about a 0.045 probability of getting a sample proportion of 0.56 or higher in a random sample of 200 students just by chance.
- Example 2Calculator allowed
p-value from a simulation
A spinner is supposed to land on red 25% of the time. In 80 spins it lands on red 30 times (p̂ = 0.375). To test H₀: p = 0.25 vs. Hₐ: p ≠ 0.25, a student simulates 1,000 sets of 80 spins assuming p = 0.25. The simulated p̂ values are centered at 0.25; 9 of them are 0.375 or higher and 6 are 0.125 or lower. Estimate the p-value.
Show the solutionHide the solution
- Step 1: The alternative is two-sided, so count results at least as far from 0.25 as 0.375 in either direction.
- Step 2: 0.375 is 0.125 above 0.25, so the other tail is 0.125 or lower.
- Step 3: p-value ≈ (9 + 6)/1,000 = 0.015.
Answer: About 0.015. Results this far from 25% happened in only 1.5% of simulations, so the spinner appears not to land on red 25% of the time.
- Example 3Calculator allowed
Trap: the p-value as P(H₀ is true)
A student says, "The p-value is 0.20, so there's a 20% chance the null hypothesis is true." Correct the statement.
Show the solutionHide the solution
- Step 1: The p-value is calculated assuming H₀ is true, so it can't be the probability that H₀ is true.
- Step 2: It's the probability of data at least this extreme if H₀ were true.
- Step 3: A p-value of 0.20 means results like this are fairly common under H₀, so there isn't convincing evidence against it.
Answer: The p-value is P(a result this extreme or more | H₀ true) = 0.20. It's not the probability that H₀ is true. Here it means the data don't give convincing evidence against H₀.
Common mistakes
- Calling the p-value the probability that H₀ (or Hₐ) is true.
- Leaving "assuming H₀ is true" out of the interpretation.
- Using one tail for a two-sided test, or the wrong tail for a one-sided test.
- Saying a large p-value proves H₀.
On the exam
- "Interpret the p-value in context" appears often. Include: assuming H₀ (with the value and context), the probability, the observed statistic and "or more extreme."
- Simulation-based p-values show up in both sections. Count dots at least as extreme as the observed value, in the direction of Hₐ.
Connected topics
Videos
Check yourself
4 questions on 3.6 p-Values. Pick an answer to see if you got it, and why.
A student suspects that a six-sided die rolls a 6 more often than 1/6 of the time. In 60 rolls, she gets 16 sixes. She simulates 1,000 sets of 60 rolls of a fair die and finds that 52 of them have 16 or more sixes. What is the estimated p-value?
A significance test gives a p-value of 0.42. Which conclusion is appropriate?
In a test of H₀: p = 0.25 versus Hₐ: p > 0.25, the p-value is 0.003. Which statement is correct?
A city's mayor claims that 60% of the city's voters approve of the job she is doing. A reporter suspects the true percent is lower. In a random sample of 400 city voters, 222 (55.5%) say they approve.
Described scenario with invented results
Which is the correct interpretation of the p-value of about 0.033?
0 of 4 answered