AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/3)
Unit 3
15–25% of examInference for Categorical Data: Proportions
Now you use sample data to draw conclusions about whole populations. You'll build confidence intervals to estimate a population proportion or a difference between two proportions, run significance tests to weigh claims about them, and use chi-square tests to look for differences or associations in two-way tables. Writing each step clearly and in context matters as much as the arithmetic.
Study this unit
Flashcards (34)Practice questions (57)Statistics must-know sheetFree-response questions on this unit
Write your own answer, then score it with the rubric or with AI.
- Question 1: Formulating questions and collecting dataReaction time with each hand10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataA salad bar survey10 points · about 22 minutes
- Question 3: InferenceFree breakfast program (test for a proportion)10 points · about 22 minutes
- Question 3: InferenceText or email reminders (interval for two proportions)10 points · about 22 minutes
- Question 3: InferenceHow students get to school (chi-square test)10 points · about 22 minutes
- Question 4: Multi-focus questionTaco truck orders10 points · about 22 minutes
- Question 4: Multi-focus questionTesting a practice app by simulation10 points · about 22 minutes
- Question 4: Multi-focus questionPizza delivery times10 points · about 22 minutes
- Question 4: Multi-focus questionStrep tests at a school clinic10 points · about 22 minutes
Big ideas
- A sample proportion p̂ is used to estimate the population proportion p
- Before any inference, check the conditions: randomization, 10% and normality (expected counts for chi-square)
- A confidence interval gives a range of plausible values for a parameter
- A small p-value means your result would be surprising if the null hypothesis were true
- Every test risks a Type I or a Type II error, and power is the chance of catching a real effect
Full unit reviews
Longer videos that cover the whole unit. Good for a first pass or a final review.
Topics
- 3.1: Estimators
- 3.2: Sampling Distributions for Sample Proportions
- 3.3: Constructing a Confidence Interval for a Population Proportion
- 3.4: Justifying a Claim Based on a Confidence Interval for a Population Proportion
- 3.5: Setting Up a Test for a Population Proportion
- 3.6: p-Values
- 3.7: Carrying Out a Test for a Population Proportion
- 3.8: Potential Errors When Performing Tests
- 3.9: Sampling Distributions for the Difference Between Sample Proportions
- 3.10: Constructing a Confidence Interval for the Difference Between Two Population Proportions
- 3.11: Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions
- 3.12: Setting Up a Test for the Difference Between Two Population Proportions
- 3.13: Carrying Out a Test for the Difference Between Two Population Proportions
- 3.14: Setting Up a Chi-Square Test for Homogeneity or Independence
- 3.15: Carrying Out a Chi-Square Test for Homogeneity or Independence
Estimators
An estimator is a statistic you use to estimate a parameter; for example, the sample proportion p̂ estimates the population proportion p, and the value you get from one sample is called a point estimate. An estimator is unbiased when, on average, it neither overestimates nor underestimates the parameter, meaning its sampling distribution is centered on the true value.
Key terms
- estimator
- point estimate
- unbiased estimator
- biased estimator
A few quick questions on this topic, with the answers explained.
When the sample is random and the observations are independent, the sampling distribution of p̂ has mean p and standard deviation √(p(1 − p)/n). It is approximately normal when np ≥ 10 and n(1 − p) ≥ 10, and the 10% condition (sample no more than 10% of the population) matters when you sample without replacement.
Key terms
- sampling distribution of p̂
- standard deviation √(p(1 − p)/n)
- randomization condition
- 10% condition
- normality (large counts) condition
A few quick questions on this topic, with the answers explained.
A one-sample z-interval estimates a population proportion with p̂ ± z* · √(p̂(1 − p̂)/n). The part after ± is the margin of error: the critical value z* multiplied by the standard error (SE). Check the randomization, 10% and normality conditions (at least 10 observed successes and 10 observed failures), and rearrange the margin-of-error formula to find the sample size you need, using p̂ = 0.5 if you have no estimate.
Key terms
- confidence interval
- one-sample z-interval
- critical value z*
- standard error (SE)
- margin of error
- sample size for a margin of error
A few quick questions on this topic, with the answers explained.
Any one interval either captures the true proportion p or misses it, so you say you're C% confident the interval captures p, and the confidence level means that about C% of intervals built the same way from repeated random samples would capture p. Values inside the interval are plausible, which lets you judge claims, and a higher confidence level makes the interval wider while a larger sample makes it narrower.
Key terms
- interpreting a confidence interval
- confidence level
- plausible values
- interval width
A few quick questions on this topic, with the answers explained.
A significance test begins with a null hypothesis, H₀: p = p₀ (the "nothing new" claim), and an alternative hypothesis, Hₐ, that p is less than, greater than or not equal to p₀, with the parameter defined in context. The conditions match the interval's, except the normality check uses the null value: np₀ ≥ 10 and n(1 − p₀) ≥ 10.
Key terms
- null hypothesis (H₀)
- alternative hypothesis (Hₐ)
- one-sided vs. two-sided test
- one-sample z-test
- conditions for a test
A few quick questions on this topic, with the answers explained.
p-Values
The p-value is the probability, assuming H₀ is true, of getting a test statistic as extreme as the one you observed, or more extreme, in the direction of Hₐ; with a simulation, it's the share of simulated results that are at least that extreme. The smaller the p-value, the more convincing the evidence for Hₐ, but a large p-value never proves that H₀ is true.
Key terms
- p-value
- null distribution
- test statistic
- convincing evidence
A few quick questions on this topic, with the answers explained.
The test statistic is z = (p̂ − p₀) ÷ √(p₀(1 − p₀)/n), and you get its p-value from the standard normal distribution. If the p-value ≤ α (the significance level), reject H₀ and say there is convincing evidence for Hₐ; otherwise, fail to reject H₀, and either way state the conclusion in context without claiming certainty.
Key terms
- z test statistic
- significance level (α)
- reject / fail to reject H₀
- conclusion in context
A few quick questions on this topic, with the answers explained.
A Type I error is rejecting a null hypothesis that is actually true, and a Type II error is failing to reject a null hypothesis that is actually false. The chance of a Type I error is α, the chance of a Type II error is 1 − power, and power (the chance of correctly rejecting a false H₀) goes up with a larger sample, a smaller standard error, a larger α or a true value farther from the null value.
Key terms
- Type I error
- Type II error
- power
- significance level (α)
A few quick questions on this topic, with the answers explained.
For two independent random samples, or a randomized experiment, the sampling distribution of p̂₁ − p̂₂ has mean p₁ − p₂ and standard deviation √(p₁(1 − p₁)/n₁ + p₂(1 − p₂)/n₂). It is approximately normal when each sample has at least 10 expected successes and 10 expected failures.
Key terms
- difference in sample proportions (p̂₁ − p̂₂)
- independent samples
- standard deviation of p̂₁ − p̂₂
- normality condition
A few quick questions on this topic, with the answers explained.
A two-sample z-interval estimates p₁ − p₂ with (p̂₁ − p̂₂) ± z* · √(p̂₁(1 − p̂₁)/n₁ + p̂₂(1 − p̂₂)/n₂). The data need to come from two independent random samples or a randomized experiment, and each sample needs at least 10 observed successes and 10 observed failures.
Key terms
- two-sample z-interval
- point estimate p̂₁ − p̂₂
- standard error
- margin of error
A few quick questions on this topic, with the answers explained.
Interpret the interval as a range of plausible values for the difference p₁ − p₂, in context. If every value in the interval is above 0 (or every value is below 0), that's convincing evidence the proportions differ; if 0 is inside the interval, "no difference" is still plausible.
Key terms
- interpreting the interval
- confidence level
- 0 inside the interval
- plausible values
A few quick questions on this topic, with the answers explained.
A two-sample z-test checks whether two population proportions differ, starting from H₀: p₁ = p₂ (the same as p₁ − p₂ = 0). For the normality condition you use the combined (pooled) proportion p̂c, the total successes divided by the total of both sample sizes, and check that n₁p̂c, n₁(1 − p̂c), n₂p̂c and n₂(1 − p̂c) are all at least 10.
Key terms
- two-sample z-test
- H₀: p₁ = p₂
- combined (pooled) proportion p̂c
- conditions for a test
A few quick questions on this topic, with the answers explained.
The test statistic is z = (p̂₁ − p̂₂ − 0) ÷ √(p̂c(1 − p̂c)(1/n₁ + 1/n₂)), and the p-value comes from the standard normal distribution. Compare the p-value with α, then write a conclusion about the difference between the two populations in context.
Key terms
- z test statistic
- p-value
- significance level (α)
- conclusion in context
A few quick questions on this topic, with the answers explained.
A chi-square (χ²) statistic measures how far the observed counts are from the counts you'd expect if H₀ were true; χ² distributions take only positive values, are skewed right and get less skewed as the degrees of freedom increase. Use a test for homogeneity to compare one categorical variable's distribution across several populations or treatments, and a test for independence to look for an association between two categorical variables in one population, checking that every expected count is greater than 5.
Key terms
- chi-square distribution
- degrees of freedom (df)
- test for homogeneity
- test for independence
- expected counts condition
A few quick questions on this topic, with the answers explained.
Each expected count is (row total × column total) ÷ table total, and the test statistic χ² = Σ (observed − expected)² / expected adds up over every cell of the table. Find the p-value from a χ² distribution with df = (number of rows − 1)(number of columns − 1), then compare it with α and state your conclusion in context.
Key terms
- expected count
- chi-square statistic (χ²)
- df = (rows − 1)(columns − 1)
- p-value
- conclusion in context
A few quick questions on this topic, with the answers explained.