AP® Psychology review sheet from Aim for Five (aimforfive.com/psych/units/3/3-8)
Unit 3 · Topic 3.8
3.8 Operant Conditioning
In operant conditioning, the consequences of a behavior change how likely you are to repeat it: reinforcement makes a behavior more likely, and punishment makes it less likely. 'Positive' means something is added and 'negative' means something is taken away. Shaping, the type of reinforcer and the schedule of reinforcement all affect how fast a behavior is learned and how long it lasts.
Key terms
- positive and negative reinforcement
- positive and negative punishment
- shaping
- reinforcement schedules
- primary and secondary reinforcers
- learned helplessness
The law of effect and the four consequences
Edward Thorndike watched cats escape puzzle boxes and proposed the law of effect: behaviors followed by satisfying consequences become more likely, and behaviors followed by unpleasant consequences become less likely. B. F. Skinner built on this, studying animals in an operant chamber (a 'Skinner box') where pressing a lever could earn food or stop a shock.
Reinforcement increases a behavior; punishment decreases it. 'Positive' and 'negative' don't mean good or bad. Positive means a stimulus is added, and negative means a stimulus is removed.
| Consequence | What happens | Effect on behavior | Example |
|---|---|---|---|
| Positive reinforcement | Add something pleasant | Increases | Getting paid for mowing the lawn |
| Negative reinforcement | Remove something unpleasant | Increases | Buckling your seatbelt stops the car's beeping |
| Positive punishment | Add something unpleasant | Decreases | Getting a speeding ticket |
| Negative punishment | Remove something pleasant | Decreases | Losing your phone for breaking curfew |
Reinforcers, shaping and limits
Primary reinforcers satisfy biological needs, like food, water and warmth. Secondary (conditioned) reinforcers gain their power by being linked with primary ones, like money, grades or praise.
Generalization and discrimination happen in operant conditioning too. A child rewarded for saying 'please' to a parent may say it to teachers too (generalization), and a dog may learn that sitting earns treats only when a certain person gives the command (discrimination).
Shaping teaches complex behaviors by reinforcing successive approximations, steps that get closer and closer to the goal. To teach a dog to roll over, you first reward lying down, then lying on its side, then rolling partway, then the full roll. But there are limits: animals trained this way sometimes drift back toward instinctive behaviors, called instinctive drift. Raccoons trained to drop coins in a box kept rubbing the coins together instead, as they would with food.
Superstition and learned helplessness
Superstitious behavior forms when a reinforcer happens to follow a behavior that didn't actually cause it. Skinner gave pigeons food at regular times no matter what they did, and many repeated whatever they happened to be doing, like turning in circles. A player who wore certain socks during a win and keeps wearing them is showing the same thing.
Learned helplessness happens when an organism learns that it has no control over unpleasant events, so it stops trying even when escape becomes possible. Martin Seligman found that dogs exposed to shocks they couldn't escape later failed to escape shocks they easily could have. Learned helplessness has been linked to depression (5.4).
Schedules of reinforcement
Continuous reinforcement rewards every correct response. It leads to fast learning but also fast extinction once rewards stop. Partial (intermittent) reinforcement rewards only some responses; learning is slower, but the behavior resists extinction much more. Partial schedules are based on either the number of responses (ratio) or the time passed (interval), on a fixed or variable basis.
| Schedule | Rule | Response pattern | Example |
|---|---|---|---|
| Fixed-ratio | After a set number of responses | High rate, brief pause after each reward | A free coffee after every 10 purchases |
| Variable-ratio | After an unpredictable number of responses | Highest, steady rate; most resistant to extinction | Slot machines; fishing |
| Fixed-interval | First response after a set amount of time | Slow after each reward, speeding up as the time nears (scalloped graph) | Checking the oven near the end of baking time; studying more as a scheduled test approaches |
| Variable-interval | First response after an unpredictable amount of time | Slow, steady rate | Checking for texts; pop quizzes |
Worked examples
Try each one yourself first, then open the solution.
- Example 1
Negative reinforcement or punishment?
Classify each. (a) Jada takes aspirin, and her headache goes away; she takes aspirin more often for headaches. (b) A teen loses driving privileges for texting while driving, and texts less. (c) A child gets extra chores for lying, and lies less.
Show the solutionHide the solution
- Step 1: Ask two questions each time: did the behavior increase or decrease? Was something added or removed?
- Step 2: (a) Taking aspirin increases, so it's reinforcement. The headache (unpleasant) is removed, so it's negative reinforcement.
- Step 3: (b) Texting decreases, so it's punishment. Driving privileges (pleasant) are removed, so it's negative punishment.
- Step 4: (c) Lying decreases, so it's punishment. Chores (unpleasant) are added, so it's positive punishment.
Answer: (a) Negative reinforcement, (b) negative punishment, (c) positive punishment.
- Example 2
Reading a schedule from data
Two groups of rats press a lever for food. Group A is rewarded every 5th press; Group B is rewarded after an unpredictable number of presses averaging 5. When food stops completely, Group A stops pressing after about 40 presses, and Group B keeps pressing for about 160. Name the schedules and explain the difference.
Show the solutionHide the solution
- Step 1: Group A is rewarded after a set number of responses: fixed-ratio. Group B is rewarded after a varying number: variable-ratio.
- Step 2: Compare persistence: 160 ÷ 40 = 4, so Group B made four times as many presses during extinction.
- Step 3: With a variable schedule, the rats never know when the next reward is coming, so a long run without food doesn't signal that rewards have stopped.
- Step 4: This is why variable-ratio schedules are the most resistant to extinction.
Answer: Group A: fixed-ratio; Group B: variable-ratio. The unpredictable variable-ratio schedule made the behavior about four times more resistant to extinction.
Common mistakes
- Treating negative reinforcement as punishment. Negative reinforcement removes something unpleasant and increases behavior; punishment always decreases behavior.
- Reading 'positive' as 'good.' Positive means something is added, whether pleasant (reward) or unpleasant (punishment).
- Mixing up ratio and interval schedules. Ratio is about the number of responses; interval is about the time that has passed.
- Thinking continuous reinforcement creates the most lasting behavior. It's fast to learn but fast to extinguish; variable-ratio lasts longest.
On the exam
- Expect many scenarios that ask you to classify a consequence or a schedule, so always ask: increase or decrease? Added or removed? Number or time? Fixed or variable?
- Graph questions may describe cumulative response lines; know that fixed-interval produces a scalloped pattern and variable-ratio a steep, steady slope.
Connected topics
Videos
Check yourself
4 questions on 3.8 Operant Conditioning. Pick an answer to see if you got it, and why.
| Reinforcement schedule | Mean presses per minute during training | Mean presses during a 1-hour session with no reinforcement |
|---|---|---|
| Fixed-ratio (every 10th press) | 52 | 310 |
| Variable-ratio (an average of every 10th press) | 61 | 920 |
| Fixed-interval (first press after each 1 minute) | 14 | 180 |
| Variable-interval (first press after an average of 1 minute) | 22 | 540 |
Hypothetical data from rats trained to press a lever for food
Which schedule is most similar to the payout pattern of a slot machine, and what do the data show about it?
Which conclusion is best supported by the data?
When 16-year-old Eli comes home an hour after curfew, his parents take away his phone for the weekend, and he starts coming home on time. Taking away the phone is an example of
A car beeps loudly until the driver fastens her seat belt. Over time, the driver starts buckling up as soon as she gets in. The driver's behavior is being shaped by
0 of 4 answered