AP® Psychology review sheet from Aim for Five (aimforfive.com/psych/units/2/2-8)
Unit 2 · Topic 2.8
2.8 Intelligence and Achievement
Psychologists still debate whether intelligence is one general ability (g) or several, and any test of it must be standardized, reliable and valid. IQ scores have risen over generations (the Flynn effect), vary more within groups than between them, and are affected by poverty, discrimination, unequal schooling and stereotype threat. Test scores have been misused to limit people's opportunities, and beliefs about whether ability can grow affect achievement.
Key terms
- general intelligence (g)
- intelligence quotient (IQ)
- reliability and validity
- Flynn effect
- stereotype threat
- growth and fixed mindset
What is intelligence?
There's no single agreed definition. One view, starting with Charles Spearman, holds that a general intelligence factor, called g, underlies performance on many different tasks, since people who do well on one kind of mental test tend to do well on others. Other theorists argue intelligence is made of several distinct abilities. Howard Gardner proposed multiple intelligences, such as linguistic, logical-mathematical, musical and interpersonal. Robert Sternberg proposed three: analytical, creative and practical.
Because definitions are debated, and because the people who design tests bring their own cultural assumptions, measuring intelligence can be biased.
How intelligence is measured
Early tests, starting with Alfred Binet's work in France in the early 1900s, estimated a child's mental age: the age level at which they performed. The original intelligence quotient (IQ) was mental age divided by chronological age, times 100. A 10-year-old performing like a typical 12-year-old scored 12 ÷ 10 × 100 = 120. This formula breaks down for adults, so modern tests compare you with others your age instead. Scores are set so the average is 100, with most people falling between 85 and 115. Schools today often use IQ scores as one piece of information when deciding which students qualify for special educational services.
A useful test must follow sound psychometric principles (the science of measuring mental traits):
- Standardized: given and scored with the same procedures and conditions for everyone, and compared with a large, representative group that took the test earlier.
- Reliable: gives consistent results. Test-retest reliability means similar scores when the same person retakes the test. Split-half reliability means scores on two halves of the test (like odd versus even questions) agree.
- Valid: measures what it claims to measure. Construct validity means it really measures the trait (the 'construct') it's meant to. Predictive validity means scores predict the future performance they're supposed to, like grades.
Fairness, culture and context
Stereotype threat is the worry that you might confirm a negative stereotype about your group, which can lower performance. Stereotype lift is a boost in performance when people are aware of a negative stereotype about another group they're compared with. Researchers try to build socioculturally responsive tests that reduce both.
The Flynn effect is the rise in average IQ scores across much of the world over the 20th century. The rise was too fast to be genetic, so it points to social factors such as better nutrition, health care, education and living standards.
IQ scores vary more within any group than between groups, so a group average tells you very little about any individual. Poverty, discrimination and educational inequities can lower scores for individuals and groups, and personal and cultural biases can affect how scores are interpreted.
Misuse of testing, and achievement
Intelligence scores have been used to limit access to jobs, military ranks and schools, and in the early 1900s were used to argue for restricting immigration to the US, often based on tests that favored people who were fluent in English and familiar with American culture. This connects to the eugenics movement (1.1).
Achievement tests measure what you've already learned, like an end-of-unit exam. Aptitude tests predict how well you'll do in the future, like a test used for admission to a program.
Carol Dweck's research shows that mindset matters. People with a fixed mindset believe intelligence is set at birth, so they may avoid challenges. People with a growth mindset believe ability grows with effort and good strategies, so they tend to persist after setbacks. You won't be tested on labels for specific cognitive abilities or disabilities.
Worked examples
Try each one yourself first, then open the solution.
- Example 1
Ratio IQ and its trap
(a) A 10-year-old scores like an average 12-year-old. What is her ratio IQ? (b) A 40-year-old scores like an average 30-year-old. Why does the ratio formula fail for him?
Show the solutionHide the solution
- Step 1: (a) Ratio IQ = mental age ÷ chronological age × 100 = 12 ÷ 10 × 100 = 120.
- Step 2: (b) Plugging in gives 30 ÷ 40 × 100 = 75, which wrongly suggests he is far below average.
- Step 3: The trap: mental abilities don't keep rising year by year in adulthood, so 'mental age' stops making sense after the teen years.
- Step 4: Modern tests solve this by comparing a person's score with others of the same age, with 100 set as the average.
Answer: (a) 120. (b) The formula assumes abilities keep growing with age, which isn't true for adults, so modern tests compare scores with same-age peers.
- Example 2
Reliable but not valid
A company measures job applicants' 'intelligence' by the circumference of their heads. Each applicant gets the same measurement every time they're measured. Is this measure reliable? Is it valid? Explain.
Show the solutionHide the solution
- Step 1: Reliability means consistency. Head measurements are the same each time, so the measure is reliable.
- Step 2: Validity means measuring what it claims to measure. Head size doesn't measure intelligence or predict job performance, so it lacks construct and predictive validity.
- Step 3: Takeaway: a test can be reliable without being valid, but it can't be valid if it isn't reliable.
Answer: Reliable (consistent results) but not valid (it doesn't measure intelligence or predict job performance).
Common mistakes
- Swapping reliability and validity. Reliability is consistency; validity is accuracy, meaning it measures what it claims to.
- Explaining the Flynn effect with genetics. Scores rose too quickly for that; social factors such as nutrition, health and schooling are the explanation.
- Using group differences in average scores to judge an individual. Scores vary more within groups than between them.
- Mixing up achievement and aptitude tests. Achievement measures what you've learned; aptitude predicts future performance.
On the exam
- Expect scenarios asking whether a test is standardized, reliable or valid, and which type of reliability or validity is described.
- For questions about fairness, use stereotype threat, the Flynn effect or within-group variation, and tie them to the specific evidence given.
Connected topics
Videos
Check yourself
4 questions on 2.8 Intelligence and Achievement. Pick an answer to see if you got it, and why.
A company designed a new test to predict how well job applicants will perform as sales associates. When 300 applicants took the test twice, two weeks apart, the correlation between their first and second scores was +0.91. However, the correlation between applicants' test scores and their supervisors' ratings of their job performance one year later was +0.08.
Scores on the test are standardized so that they form a normal distribution with a mean of 100 and a standard deviation of 15.
Hypothetical data
Based on these correlations, the test is best described as
An applicant scores 130 on the test. Approximately what percentile does this score represent?
Researchers randomly assign female college students with strong math backgrounds to two groups. One group is told that a difficult math test has shown gender differences in the past; the other is told it has shown no gender differences. The first group scores lower. This result is best explained by
A testing company finds that people today answer more questions correctly on an intelligence test than people of the same age did when the test was first normed decades ago. As a result, the company raises the number of correct answers needed for a score of 100. This change is a response to
0 of 4 answered