Skip to main content

Must-know sheet

Statistics must-know sheet

The formulas, rules, conditions and answer wording you should know cold for AP Statistics, organized around the current five-unit course. The real exam gives you a formula sheet and z, t and χ² tables on both sections, and expects a graphing calculator with statistical features (Desmos is built into Bluebook), so this sheet focuses on when to use each formula, which conditions to check and how to word your answers.

Showing all 13 sections.

Describing one quantitative variable

Unit 1

Describe a distribution: shape, center, variability, unusual features
Always cover all four (unusual features are outliers, gaps and clusters), using the variable's name and units. When comparing groups, use comparison words ("greater than," "more variable than"); two separate lists of numbers don't count as a comparison.
Shapes
Skewed right means a long tail toward the high values, and skewed left means a long tail toward the low values. Other shape words: roughly symmetric, unimodal (one peak), bimodal (two peaks) and approximately uniform (flat).
Mean: x̄ = Σxᵢ ÷ n
The balance point of the data. It gets pulled toward the tail and toward outliers, so in a right-skewed distribution the mean is usually greater than the median, and in a left-skewed one it's usually less.
Median
The middle value of the ordered data (the average of the two middle values when n is even). It is also Q2, the 50th percentile.
Sample standard deviation: s = √[Σ(xᵢ − x̄)² ÷ (n − 1)]
About how far a typical value falls from the mean. s² is the variance; s = 0 only when every value is the same. Find it with your calculator.
Range = max − min; IQR = Q3 − Q1
The IQR is the spread of the middle 50% of the data. Q1 and Q3 are the 25th and 75th percentiles.
Resistant vs. not resistant
The median and IQR barely move when there are outliers or strong skew (resistant). The mean, standard deviation and range can change a lot (not resistant). For skewed data or data with outliers, report the median and IQR.
Outlier rule 1: 1.5 × IQR
A value is an outlier if it is below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR. On the exam, show both fences and compare them with the minimum and maximum.
Outlier rule 2: 2 standard deviations
A value is an outlier if it is more than 2 standard deviations above or below the mean, that is, outside x̄ ± 2s.
Boxplot
Draws the five-number summary: min, Q1, median, Q3, max. With outliers, the whiskers stop at the most extreme values that aren't outliers, and outliers get their own marks. A boxplot can't show gaps, clusters or two peaks.
Percentile
The pth percentile is the value with p% of the data at or below it.
z-score: z = (value − mean) ÷ standard deviation
How many standard deviations a value is above (positive) or below (negative) the mean. Use it to compare values from different distributions. For sample data, use (x − x̄) ÷ s.
Changing units
Adding or subtracting a constant shifts the mean, median and quartiles by that constant but leaves the range, IQR and standard deviation unchanged. Multiplying every value by a positive constant multiplies both the center and the spread measures by that constant.
Graphs for one variable
Quantitative: dotplot, stemplot (stem-and-leaf plot), histogram (the bin width can change how it looks) and boxplot. Categorical: bar chart and pie chart, using counts or relative frequencies (proportions).

Collecting data: questions, samples and experiments

Unit 1

Investigative question
A good question has three parts: the variable(s) to collect data on; what the analysis will do (estimate a parameter, or test a claim about it in a stated direction, such as "greater than" or "associated"); and what kind of conclusion you can draw (which population, and cause and effect only if treatments are randomly assigned). Set it before collecting data and don't change it after seeing the results.
Population, sample, parameter, statistic
A parameter describes the whole population (p, μ, σ); a statistic describes a sample (p̂, x̄, s) and is used to estimate the parameter. Memory hook: Population–Parameter, Sample–Statistic.
Types of variables
Categorical variables place units in groups. Quantitative variables are counted or measured numbers: discrete (countable values) or continuous (any value in an interval).
Census
Collects data from every member of the population, so no estimating is needed, but it's usually too slow or costly.
Observational study vs. experiment
An observational study records what happens without imposing treatments, so a confounding variable might explain the results. A prospective study follows its units forward in time; a retrospective study looks back at past data; a survey asks people a set list of questions. An experiment assigns treatments to experimental units (subjects, if people).
Explanatory variable (factor), levels, treatments, response
In an experiment, the explanatory variable (factor) is what the researcher changes, and its levels are the treatments; with more than one factor, each combination of levels is a treatment. The response variable is the outcome measured on each unit afterward.
Confounding variable
A variable linked to both the explanatory variable and the response, so you can't tell which one caused the change in the response.
Simple random sample (SRS)
Every possible group of n individuals has the same chance of being chosen. Label the units, then use a random number generator to pick n different labels (ignore repeats).
Stratified random sample
Split the population into strata of similar units (for example, grade level), then take an SRS from every stratum. It works best when units within a stratum are alike, and it gives more precise estimates.
Cluster sample
Split the population into clusters (for example, homerooms), randomly choose some whole clusters and use everyone in them. Ideally each cluster is a mix that looks like the population. It's often chosen because it's convenient and cheaper.
Systematic random sample
Pick a random starting point, then take every kth unit from an ordered list.
Bias
A sampling method is biased if it tends to produce statistics that are too high, or too low, compared with the parameter. A bigger sample does not fix bias.
Sources of bias
Voluntary response (people choose to respond, usually the ones with strong opinions); convenience sampling; undercoverage (part of the population can't be chosen); nonresponse (chosen people don't answer); response bias (wording, interviewer effects or lying). Say which direction the bias pushes the estimate.
Principles of a well-designed experiment
Compare at least two treatments (one may be a control); randomly assign the treatments; replicate (use more than one unit in each treatment); and control other variables that could affect the response. Random assignment makes the groups roughly alike on extraneous variables, so a difference in response can be credited to the treatments.
Completely randomized design
All experimental units are randomly assigned to the treatments, with no blocking.
Randomized block design
Group units into blocks that are alike on a variable expected to affect the response, then randomly assign treatments within each block. It reduces variability by comparing like with like.
Matched pairs design
A block of two: either pairs of very similar units, each getting a different treatment, or each unit getting both treatments in a random order. Analyze the differences.
Control group, placebo, blinding
A control group gives a baseline to compare with. A placebo is a fake treatment, used because of the placebo effect. Single-blind: either the subjects or the researchers who work with them don't know who gets which treatment; double-blind: neither the subjects nor those researchers know.
Scope of inference
Random sampling lets you generalize to the population the sample came from; without it (volunteers, convenience), generalize only to units like those in the study. Random assignment of treatments lets you conclude cause and effect. Without random assignment, you can show an association, not causation.

Two categorical variables and probability rules

Unit 2

Joint, marginal and conditional relative frequency
Joint: a cell ÷ the table total. Marginal: a row or column total ÷ the table total. Conditional: a cell ÷ its own row or column total.
Association between two categorical variables
The variables are associated if the conditional distributions differ from group to group, for example in a segmented bar chart or mosaic plot. If they're about the same, there's little or no association.
Simulation
Use a chance device (random numbers) to imitate a random process many times. The estimated probability is the share of trials in which the event happened.
Law of large numbers
Over many independent trials, the relative frequency of an event settles closer and closer to its true probability. It says nothing about the short run.
Basic rules
0 ≤ P(A) ≤ 1, and the probabilities of all outcomes in the sample space add up to 1. With equally likely outcomes, P(A) = (outcomes in A) ÷ (total outcomes).
Complement rule: P(not A) = 1 − P(A)
Handy for "at least one" questions: P(at least one) = 1 − P(none).
Mutually exclusive (disjoint) events
A and B can't both happen, so P(A ∩ B) = 0 and P(A ∪ B) = P(A) + P(B). Two disjoint events that each have a positive probability are never independent.
General addition rule: P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
"A or B" means A, B or both. Subtract the overlap so it isn't counted twice.
Conditional probability: P(A | B) = P(A ∩ B) ÷ P(B)
The chance of A given that B happened. In a two-way table, look only at the row or column for B.
General multiplication rule: P(A ∩ B) = P(A) · P(B | A)
Use for "A and B," especially in tree diagrams: multiply along the branches, then add the branch results that lead to the event you want.
Independent events
A and B are independent if P(A | B) = P(A), meaning knowing B happened doesn't change the chance of A. Then P(A ∩ B) = P(A) · P(B). To justify independence, show a calculation, such as comparing P(A | B) with P(A).
Notation
∩ means "and" (both), ∪ means "or" (at least one), and | means "given that." P(A | B) and P(B | A) are usually different.

Random variables, binomial and normal distributions

Unit 2

Discrete probability distribution
Lists every value of X with its probability. Each probability is between 0 and 1, and they add up to 1. A cumulative distribution gives P(X ≤ x).
Mean (expected value): μX = Σ xᵢ · P(xᵢ)
The long-run average value of X over many, many repetitions. It doesn't have to be a possible value of X.
Standard deviation: σX = √[Σ (xᵢ − μX)² · P(xᵢ)]
How far values of X typically fall from μX over many repetitions.
Binomial setting (check all four)
Two outcomes on each trial (success or failure); trials independent; a fixed number of trials, n; the same probability of success, p, on every trial. X counts the successes.
Binomial probability: P(X = x) = C(n, x) · pˣ · (1 − p)ⁿ⁻ˣ
C(n, x) = n! ÷ [x!(n − x)!] counts the arrangements, for x = 0, 1, …, n. Technology works too (binompdf gives P(X = x) and binomcdf gives P(X ≤ x) on a TI-84), but on free response also name the distribution and its values, such as "binomial, n = 20, p = 0.3, P(X ≥ 5)."
Binomial "at least" and "more than"
P(X ≥ k) = 1 − P(X ≤ k − 1), and P(X > k) = 1 − P(X ≤ k). Write which values are included before you calculate.
Binomial mean and SD: μX = np, σX = √[np(1 − p)]
Interpret the mean as the average number of successes over many sets of n trials.
Normal distribution
A symmetric, unimodal, bell-shaped curve centered at μ with spread σ. The total area is 1, and the area over an interval is a probability. Because it's continuous, P(X = a) = 0, so < and ≤ give the same answer.
Empirical (68–95–99.7) rule
In a normal distribution, about 68% of values are within 1σ of μ, 95% within 2σ and 99.7% within 3σ. Use it only when the distribution is approximately normal.
Normal calculations
Standardize with z = (x − μ) ÷ σ; the standard normal distribution has mean 0 and SD 1. The z-table gives the area to the left of z, and the area to the right is 1 − (area to the left). To work backward from a percentile, use x = μ + zσ, and always show the distribution, its mean and SD, the boundary and the area you want.
Common critical values z*
90% → 1.645, 95% → 1.960, 98% → 2.326, 99% → 2.576. z* cuts off the middle C% of the standard normal curve.

Sampling distributions

Units 2, 3, 4

Sampling distribution
The distribution of a statistic (such as p̂ or x̄) over all possible samples of the same size from the same population. You can approximate one by simulating many samples.
Randomization distribution
The values of a statistic you get by repeatedly reshuffling the response values at random between treatment groups. It shows what results look like when the treatment has no effect.
Unbiased estimator
A statistic whose sampling distribution is centered at the parameter it estimates, as p̂ is for p and x̄ is for μ. Bias is about the center; variability is about the spread.
Sample proportion p̂
Mean p, standard deviation √[p(1 − p) ÷ n]. Approximately normal when np ≥ 10 and n(1 − p) ≥ 10.
Difference of sample proportions p̂₁ − p̂₂
Mean p₁ − p₂, standard deviation √[p₁(1 − p₁) ÷ n₁ + p₂(1 − p₂) ÷ n₂]. Approximately normal when n₁p₁, n₁(1 − p₁), n₂p₂ and n₂(1 − p₂) are all at least 10.
Sample mean x̄
Mean μ, standard deviation σ ÷ √n. Exactly normal when the population is normal; approximately normal when n ≥ 30 by the central limit theorem.
Difference of sample means x̄₁ − x̄₂
Mean μ₁ − μ₂, standard deviation √(σ₁² ÷ n₁ + σ₂² ÷ n₂). Normal when both populations are normal; approximately normal when n₁ ≥ 30 and n₂ ≥ 30.
Central limit theorem (CLT)
For a large enough random sample, the sampling distribution of x̄ is approximately normal even when the population isn't, and the bigger n is, the closer to normal it gets. A strongly skewed population may need far more than 30.
Bigger samples, less variability
The standard deviation of p̂ or x̄ shrinks with √n: four times the sample size cuts it in half. Sample size changes the spread, not the center.
Standard deviation vs. standard error (SE)
A standard deviation uses parameters (p, σ). The standard error uses sample values (p̂, s) in the same formula, because the parameters are unknown.
10% condition for these formulas
When sampling without replacement, the standard deviation formulas are only accurate if n ≤ 10% of the population size.

Inference for proportions: formulas

Unit 3

Every confidence interval: point estimate ± margin of error
Margin of error = critical value × standard error. The margin of error is half the interval's width, and the point estimate sits in the middle.
Every test statistic: (statistic − null value) ÷ standard error
How many standard errors the sample result is from what H₀ predicts.
One-sample z-interval for p: p̂ ± z* · √[p̂(1 − p̂) ÷ n]
Estimates a single population proportion. The SE uses p̂ because p is unknown.
Sample size for a margin of error: n = (z* ÷ ME)² · p̂(1 − p̂)
Always round up to the next whole number. If you have no estimate of p, use p̂ = 0.5, which gives the largest (safest) n. Example: ME = 0.03 at 95% with p̂ = 0.5 gives 1,067.1, so n = 1,068.
One-sample z-test for p: z = (p̂ − p₀) ÷ √[p₀(1 − p₀) ÷ n]
H₀: p = p₀. The SE uses the null value p₀, because the test assumes H₀ is true. Get the p-value from the standard normal distribution.
Two-sample z-interval for p₁ − p₂
(p̂₁ − p̂₂) ± z* · √[p̂₁(1 − p̂₁) ÷ n₁ + p̂₂(1 − p̂₂) ÷ n₂]. Define which group is 1 and which is 2, and keep that order.
Two-sample z-test for p₁ − p₂
H₀: p₁ = p₂ (p₁ − p₂ = 0). z = (p̂₁ − p̂₂ − 0) ÷ √[p̂c(1 − p̂c)(1 ÷ n₁ + 1 ÷ n₂)], using the combined (pooled) proportion because H₀ says the two proportions are equal.
Combined (pooled) proportion: p̂c = (x₁ + x₂) ÷ (n₁ + n₂)
Total successes ÷ total sample size, the same as (n₁p̂₁ + n₂p̂₂) ÷ (n₁ + n₂). Use it only in the two-proportion test, never in the interval.
What changes an interval's width
A higher confidence level gives a bigger z*, so a wider interval. A larger sample gives a smaller SE, so a narrower interval (four times n halves the width). Neither one fixes bias from a bad sample.

Chi-square tests for two-way tables

Unit 3

Which chi-square test?
Homogeneity: independent random samples from two or more populations (or groups in a randomized experiment), one categorical variable, asking whether its distribution is the same in every group. Independence: one random sample, two categorical variables, asking whether they are associated in the population. Both work on a two-way table of counts.
Hypotheses: homogeneity
H₀: the distribution of [variable] is the same for all [populations or treatments]. Hₐ: the distribution of [variable] is not the same for all of them. Use the context's words.
Hypotheses: independence
H₀: there is no association between [variable 1] and [variable 2] in [population] (they are independent). Hₐ: there is an association (they are not independent).
Expected count = (row total × column total) ÷ table total
The count you'd expect in a cell if H₀ were true. Expected counts don't have to be whole numbers.
χ² = Σ (observed − expected)² ÷ expected
Add this up over every cell, using counts, never proportions. A bigger χ² means the observed counts are farther from what H₀ predicts.
Degrees of freedom: df = (rows − 1)(columns − 1)
Count only the categories, not the total row or column.
χ² distributions
They take only values ≥ 0, are skewed right, have mean equal to df and get less skewed as df grows. The p-value is always the area to the right of the χ² value.
Where the difference is
If you reject H₀, the cells with the largest (observed − expected)² ÷ expected terms contribute most to the difference; compare their observed and expected counts.

Inference for means: formulas

Unit 4

Why t instead of z
When σ is unknown, use s, and the standardized statistic follows a t-distribution. t-distributions are symmetric and bell-shaped with heavier tails than the standard normal, so t* is bigger than z*. As df grows, t gets closer to the standard normal.
One-sample t-interval for μ: x̄ ± t* · s ÷ √n
df = n − 1. Get t* from the t-table or your calculator.
One-sample t-test for μ: t = (x̄ − μ₀) ÷ (s ÷ √n)
H₀: μ = μ₀, with df = n − 1. Get the p-value from the t-distribution.
Paired data (matched pairs): one-sample t on the differences
Take the difference within each pair (keep the same order every time), then use x̄d, sd and n = the number of pairs. The parameter is μd, the population mean difference; the usual test is H₀: μd = 0, with df = n − 1.
Two-sample t-interval for μ₁ − μ₂
(x̄₁ − x̄₂) ± t* · √(s₁² ÷ n₁ + s₂² ÷ n₂). Use technology for df; it falls between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2. Using the smaller n − 1 is the safe choice by hand.
Two-sample t-test for μ₁ − μ₂
H₀: μ₁ = μ₂ (μ₁ − μ₂ = 0). t = (x̄₁ − x̄₂ − 0) ÷ √(s₁² ÷ n₁ + s₂² ÷ n₂), with the p-value from technology. Don't pool the standard deviations.
Paired or two-sample?
Paired: the same units measured twice, or units deliberately matched in pairs, so the data come in linked pairs. Two-sample: two separate, independent groups. Using a two-sample procedure on paired data is a common, costly mistake.

Conditions to check for every procedure

Units 3, 4

Randomization condition
One-sample procedures and the chi-square test for independence: a random sample (for a one-sample mean or mean difference, a randomized experiment also works). Two-sample procedures and chi-square homogeneity: two (or more) independent random samples, or a randomized experiment.
10% condition
When sampling without replacement, each sample must be no more than 10% of its population: n ≤ 0.10N. It keeps the observations close to independent. It isn't needed for a randomized experiment, and it is not the same thing as the randomization condition.
Normality: one-proportion interval
Observed successes and failures are both at least 10: np̂ ≥ 10 and n(1 − p̂) ≥ 10.
Normality: one-proportion test
Use the null value: np₀ ≥ 10 and n(1 − p₀) ≥ 10.
Normality: two-proportion interval
Observed successes and failures in each sample are all at least 10: n₁p̂₁, n₁(1 − p̂₁), n₂p̂₂ and n₂(1 − p̂₂).
Normality: two-proportion test
Using the combined proportion, n₁p̂c, n₁(1 − p̂c), n₂p̂c and n₂(1 − p̂c) are all at least 10.
Expected counts: chi-square tests
Every expected count (not observed count) is greater than 5, as the course framework puts it (many textbooks say at least 5). Show the smallest expected count, or list them all.
Sample data condition: one-sample t (and paired t)
The population is approximately normal, or n ≥ 30, or, if n < 30, a graph of the sample shows no strong skew and no outliers. For paired data, check the differences, not the two original lists.
Sample data condition: two-sample t
Both n₁ ≥ 30 and n₂ ≥ 30, or both populations are approximately normal. If either sample is under 30, graphs of both samples should show no strong skew and no outliers.
How to write conditions
Check each one in context with numbers, for example "the 40 students were a random sample of the school," "40 ≤ 10% of 1,200," "80(0.35) = 28 ≥ 10." Writing just "SRS" or "normal" without evidence earns no credit.

Significance tests: hypotheses, p-values and errors

Units 3, 4

Hypotheses
Write them about parameters (p, μ, p₁ − p₂, μ₁ − μ₂, μd), never statistics, and define the parameter in context. H₀ has the equals sign; Hₐ uses <, > or ≠ and is chosen before you look at the data.
Significance level α
How much evidence you require, set before the test; 0.05 is the default if none is given. It is also the probability of a Type I error.
p-value
Assuming H₀ is true, the probability of getting a result as extreme as the one observed, or more extreme, in the direction of Hₐ. For ≠, add the areas in both tails (double one tail for z or t).
p-value from a simulation
When the null distribution comes from a simulation or randomization (reshuffling), the p-value is the proportion of simulated statistics at least as extreme as the observed one, in the direction of Hₐ (both tails for ≠).
Decision
If p-value ≤ α, reject H₀: there is convincing evidence for Hₐ. If p-value > α, fail to reject H₀: there is not convincing evidence for Hₐ. Always give the comparison, such as "0.012 < 0.05."
Never "accept H₀"
A large p-value means the data are consistent with H₀, not that H₀ is true. Lack of evidence for Hₐ is not evidence for H₀.
Type I error
Rejecting H₀ when H₀ is actually true: finding convincing evidence for Hₐ when Hₐ is false. P(Type I) = α.
Type II error
Failing to reject H₀ when Hₐ is actually true: missing a real effect. P(Type II) = 1 − power.
Power
The probability of correctly rejecting a false H₀. It goes up with a larger sample size (smaller standard error), a larger α, or a true parameter value farther from the null value.
Error trade-off
Lowering α makes a Type I error less likely but a Type II error more likely (lower power). Increasing the sample size lowers the chance of a Type II error without raising α. Pick α by consequences: if a Type I error is worse, use a smaller α; if a Type II error is worse, use a larger α or a bigger sample.
Intervals and two-sided tests agree
A C% confidence interval matches a two-sided test with α = 1 − C. If the null value is outside the interval, reject H₀; if it's inside, the null value is plausible. For a difference, check whether 0 is in the interval. For means the match is exact; for proportions it's almost always the same, because the test's SE uses p₀ (or p̂c) and the interval's uses p̂.
Spot the question type
"Estimate" or "construct an interval" means a confidence interval. "Is there convincing evidence…?" means a significance test, not just a description of the data.

Regression with two quantitative variables

Unit 5

Scatterplot
Put the explanatory variable on the x-axis and the response on the y-axis. Describe form (linear or curved), direction (positive or negative), strength (strong, moderate, weak) and unusual features (outliers, clusters), in context.
Correlation r
Measures the direction and strength of a linear relationship, from −1 to 1, with no units. It doesn't change if you swap x and y or change units, it isn't resistant to outliers, and it only describes linear relationships.
Correlation isn't causation
A strong r shows an association, not that x causes changes in y. A strong r also doesn't prove a line is the right model; check the residual plot.
Regression line: ŷ = a + bx
ŷ (y-hat) is the predicted response, b is the slope and a is the y-intercept. Find a, b and r with technology. On computer output, "Constant" is a and the coefficient on the x variable is b.
Least-squares regression line (LSRL)
The line that makes the sum of the squared residuals as small as possible. It always passes through (x̄, ȳ), and the residuals add up to 0.
Residual = y − ŷ (observed − predicted)
A positive residual means the line underestimated (the point is above the line); a negative residual means it overestimated (the point is below the line).
Residual plot
Residuals plotted against x (or ŷ). Random scatter with no pattern means a linear model is appropriate; a curved pattern means it isn't.
Interpolation vs. extrapolation
Predicting within the range of the x-data is interpolation. Predicting outside it is extrapolation, which is unreliable because the pattern may not continue.
Coefficient of determination r²
The proportion (percent) of the variation in y that is explained by the linear relationship with x. r has the same sign as the slope, so r = ±√r².

Interpretation templates for free-response answers

Units 1, 2, 3, 4, 5

Confidence interval
"We are C% confident that the interval from a to b captures the true [parameter: mean or proportion, for which population, in context]."
Confidence level
"If we took many random samples of this size and built a C% interval from each one, about C% of those intervals would capture the true [parameter in context]." It does not mean there's a C% chance this one interval is right.
Interpreting a p-value
"Assuming [H₀ in context] is true, there is a [p-value] probability of getting a [statistic] as extreme as [observed value] or more extreme, in the direction of Hₐ, by chance alone."
Conclusion when p-value ≤ α
"Because the p-value of [value] ≤ α = [value], we reject H₀. There is convincing evidence that [Hₐ in context]."
Conclusion when p-value > α
"Because the p-value of [value] > α = [value], we fail to reject H₀. There is not convincing evidence that [Hₐ in context]."
Slope
"For each additional 1 [unit of x], the predicted [y in context] increases (or decreases) by about [b] [units of y]." Say "predicted"; the slope describes the model, not every individual.
y-intercept
"When [x in context] is 0, the predicted [y in context] is [a]." Add a note when x = 0 is far outside the data or the prediction is impossible, since then it has no sensible meaning.
r²
"About [r² as a percent] of the variation in [y] is explained by the linear relationship with [x]."
Interpreting r
"There is a [strong/moderate/weak], [positive/negative] linear relationship between [x in context] and [y in context]." Then say what the direction means, for example "longer wolves tend to weigh more."
Residual
"The actual [y] for this [unit] was [residual] [units] above (or below) the value predicted by the regression line."
Standard deviation
"The [variable] typically varies by about [s] [units] from the mean of [x̄] [units]."
z-score
"This [value] is [z] standard deviations above (or below) the mean [variable] for [group]."
Expected value of a random variable
"Over many, many repetitions, the average [X in context] would be about [μX]."
Standard deviation of a sampling distribution
"In repeated random samples of size n, the sample [proportion or mean] typically varies by about [SD] from the true [parameter]."
Type I or II error, in context
Say what was concluded and what is actually true, then name the consequence, for example: "We conclude the new drug works better when it really doesn't, so patients switch to it for no benefit."
Stating the scope of inference
Say who you can generalize to (only with random sampling) and whether you can claim cause and effect (only with random assignment), naming the feature of the study that justifies each.

Choosing the procedure and writing a full inference answer

Units 3, 4

One categorical variable, one sample
Estimate: one-sample z-interval for p. Test a claim: one-sample z-test for p.
One categorical variable, two groups (two outcomes)
Estimate: two-sample z-interval for p₁ − p₂. Test: two-sample z-test for p₁ − p₂.
Two-way tables of counts (often more than two categories)
Several groups compared on one categorical variable: chi-square test for homogeneity. One sample, two categorical variables: chi-square test for independence. Use these for any two-way table of counts, especially when a variable has more than two categories.
One quantitative variable, one sample
One-sample t-interval or t-test for μ.
Quantitative, paired data
Paired t-interval or paired t-test for μd, done on the differences.
Quantitative, two independent groups
Two-sample t-interval or t-test for μ₁ − μ₂.
Full answer, step 1: name and set up
Name the procedure in full (or give its formula), such as "two-sample z-test for the difference between two population proportions"; short names like "two-sample z-test" or "2 z-test" don't earn credit. Define the parameter in context and, for a test, state H₀, Hₐ and α.
Full answer, step 2: conditions
Check the randomization, 10% and normality (or sample data or expected counts) conditions, each in context with numbers.
Full answer, step 3: calculate
Show the formula with numbers substituted, or name the distribution or procedure and label its inputs; bare calculator syntax may not earn credit. Report the test statistic, df (for t and χ²) and the p-value, or the interval.
Full answer, step 4: conclude
Compare the p-value with α (or use the interval), then answer the question in context with non-definitive wording such as "convincing evidence," never "proves."