AP® Statistics review sheet from Aim for Five (aimforfive.com/stats/units/1)
Unit 1
20–30% of examExploring One-Variable Data and Collecting Data
Statistics starts with a question and some data. In this unit you'll describe one variable at a time with tables, graphs and summary numbers, and you'll judge how data were collected, because the way data are gathered decides which conclusions you're allowed to draw. These skills show up all over the exam.
Study this unit
Flashcards (40)Practice questions (66)Statistics must-know sheetFree-response questions on this unit
Write your own answer, then score it with the rubric or with AI.
- Question 1: Formulating questions and collecting dataThree plans for a library survey10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataMusic and typing speed experiment10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataEnergy drinks and sleep survey10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataReaction time with each hand10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataA salad bar survey10 points · about 22 minutes
- Question 1: Formulating questions and collecting dataTree cover and park temperatures10 points · about 22 minutes
- Question 2: Analyzing data and interpreting resultsBattery life of two phone models10 points · about 22 minutes
- Question 2: Analyzing data and interpreting resultsHomework minutes10 points · about 22 minutes
- Question 2: Analyzing data and interpreting resultsMain news source by age group10 points · about 22 minutes
- Question 2: Analyzing data and interpreting resultsJuice bottle fill amounts10 points · about 22 minutes
- Question 3: InferenceFree breakfast program (test for a proportion)10 points · about 22 minutes
- Question 3: InferenceText or email reminders (interval for two proportions)10 points · about 22 minutes
- Question 3: InferenceHow students get to school (chi-square test)10 points · about 22 minutes
- Question 3: InferenceLaptop versus phone typing (paired t-test)10 points · about 22 minutes
- Question 3: InferenceRide-share wait times (two-sample t-test)10 points · about 22 minutes
- Question 3: InferenceDaily screen time (interval for a mean)10 points · about 22 minutes
- Question 4: Multi-focus questionTesting a practice app by simulation10 points · about 22 minutes
- Question 4: Multi-focus questionGrip strength training10 points · about 22 minutes
Big ideas
- A good investigative question names the variable, the population and what you want to find out
- Describe a distribution's shape, center, variability and unusual features, always in context
- The median and IQR resist outliers; the mean, standard deviation and range don't
- Random sampling is what lets you generalize from a sample to a population
- Random assignment in an experiment is what allows cause-and-effect conclusions
Full unit reviews
Longer videos that cover the whole unit. Good for a first pass or a final review.
Topics
- 1.1: Introducing Statistics: What Can We Learn from Data?
- 1.2: Variables
- 1.3: Tabular Representation and Summary Statistics for One Categorical Variable
- 1.4: Graphical Representations for One Categorical Variable
- 1.5: Graphical Representations for One Quantitative Variable
- 1.6: Descriptions for One Quantitative Variable Distributions
- 1.7: Summary Statistics for One Quantitative Variable
- 1.8: Graphical Representations of Summary Statistics for One Quantitative Variable
- 1.9: Comparisons of the Distributions for One Quantitative Variable
- 1.10: The Investigative Question Revisited and Data Collection
- 1.11: Random Sampling
- 1.12: Potential Problems with Sampling
- 1.13: Experimental Design
Statistics lets you answer an investigative question about a large group (the population) by collecting data from a smaller group (the sample), since getting data from every member is usually impossible or too costly. A good question has a clear purpose, doesn't change once you see the results, and can be answered with data you're able to collect.
Key terms
- investigative question
- population (size N)
- sample (size n)
- data set
- in context
A few quick questions on this topic, with the answers explained.
Variables
An observational unit is the item or person you collect data from, and a variable is a characteristic that can differ from unit to unit: categorical (group labels) or quantitative (counted or measured numbers, which can be discrete or continuous). A parameter summarizes a whole population, while a statistic summarizes a sample and is often used to estimate the parameter.
Key terms
- observational unit
- categorical variable
- quantitative variable
- discrete vs. continuous
- parameter
- statistic
A few quick questions on this topic, with the answers explained.
A frequency table counts how many units fall in each category of a categorical variable, and a relative frequency table shows each category's proportion of the total. Proportions, percentages and relative frequencies carry the same information, and you use them to back up claims about the variable.
Key terms
- frequency table
- relative frequency
- proportion
- percentage
A few quick questions on this topic, with the answers explained.
Bar charts and pie charts show the counts or proportions in each category of one categorical variable: a bar's height or a slice's share of the circle matches that category's share of the data. You can use tables and graphs like these to compare two or more groups on the same categorical variable.
Key terms
- bar chart (bar graph)
- pie chart
- frequency
- relative frequency
A few quick questions on this topic, with the answers explained.
Dotplots, stem-and-leaf plots (stemplots) and histograms show the distribution of a quantitative variable while keeping the values in order from smallest to largest. A histogram groups values into intervals called bins, and changing the bin width can change how the graph looks.
Key terms
- dotplot
- stem-and-leaf plot
- histogram
- bin width
- distribution
A few quick questions on this topic, with the answers explained.
To describe a quantitative distribution, discuss its shape, center, variability (spread) and any unusual features such as outliers, gaps or clusters, always in context. Common shapes are skewed right (a longer right tail), skewed left, roughly symmetric, unimodal, bimodal and approximately uniform.
Key terms
- shape, center and variability
- skewed right / skewed left
- symmetric
- unimodal / bimodal / uniform
- outlier
- gaps and clusters
A few quick questions on this topic, with the answers explained.
The mean (x̄) and median measure center; the range, the interquartile range (IQR = Q3 − Q1) and the standard deviation (s) measure variability; and percentiles and quartiles describe a value's position. One common rule calls a value an outlier if it is more than 1.5 × IQR below Q1 or above Q3 (another flags values more than 2 standard deviations from the mean). Outliers barely move the median and IQR, so those are called resistant, while the mean, range and standard deviation are not.
Key terms
- mean (x̄) and median
- standard deviation (s)
- interquartile range (IQR)
- percentile
- outlier rules (1.5 × IQR, 2 standard deviations)
- resistant measure
A few quick questions on this topic, with the answers explained.
A boxplot graphs the five-number summary (minimum, Q1, median, Q3, maximum), with the box covering the middle 50% of the data; when there are outliers, the whiskers stop at the most extreme values that aren't outliers, and the outliers get their own marks. Shape links the mean and median: in a right-skewed distribution the mean is usually greater than the median, and in a left-skewed one it is usually less.
Key terms
- five-number summary
- boxplot
- quartiles (Q1 and Q3)
- mean vs. median and skew
A few quick questions on this topic, with the answers explained.
To compare distributions, compare their shapes, centers, variability and unusual features using comparison words like "greater than," not just separate lists of numbers. A z-score, z = (x − μ)/σ, tells how many standard deviations a value sits above or below the mean, so you can compare values that come from different distributions.
Key terms
- comparing distributions
- side-by-side boxplots
- back-to-back stemplot
- z-score (standardized score)
A few quick questions on this topic, with the answers explained.
The investigative question gets more detailed here: it should point to the variables to collect, the analysis to run (such as a test or a confidence interval) and the kind of conclusion you can make. You'll also tell apart a census (data from every member), an observational study (no treatments imposed, so a confounding variable could explain the results) and an experiment (treatments assigned to experimental units), and you can only generalize to the whole population when the sample was chosen at random.
Key terms
- census
- observational study
- experiment
- explanatory and response variables
- confounding variable
- generalizing to a population
A few quick questions on this topic, with the answers explained.
Random Sampling
Random sampling uses a chance process, like a random number generator, to choose the sample. The main methods are the simple random sample (every possible sample of size n is equally likely), the stratified random sample (an SRS from each group of similar units), the cluster sample (randomly choose whole groups and use everyone in them) and the systematic sample (a random start, then every kth unit).
Key terms
- simple random sample (SRS)
- stratified random sample
- cluster sample
- systematic random sample
- sampling with / without replacement
A few quick questions on this topic, with the answers explained.
Bias means the way a sample is chosen pushes a statistic in one direction, so it keeps coming out too high or too low compared with the parameter you're trying to estimate. Watch for voluntary response bias, undercoverage, nonresponse and response bias (such as leading questions), and remember that convenience and volunteer samples aren't random, so they invite bias.
Key terms
- bias
- voluntary response bias
- undercoverage
- nonresponse bias
- response bias
- convenience sample
A few quick questions on this topic, with the answers explained.
A well-designed experiment compares at least two treatments, assigns them at random, uses replication (more than one unit per treatment) and keeps other sources of variation under control. You'll learn completely randomized, randomized block and matched pairs designs, along with control groups, placebos and blinding, and why random assignment is what allows cause-and-effect conclusions.
Key terms
- random assignment
- control group and placebo
- single-blind / double-blind
- replication
- randomized block design
- matched pairs design
A few quick questions on this topic, with the answers explained.