Skip to main content

Unit 3 · Topic 3.8

3.8 Operant Conditioning

In operant conditioning, the consequences of a behavior change how likely you are to repeat it: reinforcement makes a behavior more likely, and punishment makes it less likely. 'Positive' means something is added and 'negative' means something is taken away. Shaping, the type of reinforcer and the schedule of reinforcement all affect how fast a behavior is learned and how long it lasts.

Key terms

  • positive and negative reinforcement
  • positive and negative punishment
  • shaping
  • reinforcement schedules
  • primary and secondary reinforcers
  • learned helplessness

The law of effect and the four consequences

Edward Thorndike watched cats escape puzzle boxes and proposed the law of effect: behaviors followed by satisfying consequences become more likely, and behaviors followed by unpleasant consequences become less likely. B. F. Skinner built on this, studying animals in an operant chamber (a 'Skinner box') where pressing a lever could earn food or stop a shock.

Reinforcement increases a behavior; punishment decreases it. 'Positive' and 'negative' don't mean good or bad. Positive means a stimulus is added, and negative means a stimulus is removed.

ConsequenceWhat happensEffect on behaviorExample
Positive reinforcementAdd something pleasantIncreasesGetting paid for mowing the lawn
Negative reinforcementRemove something unpleasantIncreasesBuckling your seatbelt stops the car's beeping
Positive punishmentAdd something unpleasantDecreasesGetting a speeding ticket
Negative punishmentRemove something pleasantDecreasesLosing your phone for breaking curfew

Reinforcers, shaping and limits

Primary reinforcers satisfy biological needs, like food, water and warmth. Secondary (conditioned) reinforcers gain their power by being linked with primary ones, like money, grades or praise.

Generalization and discrimination happen in operant conditioning too. A child rewarded for saying 'please' to a parent may say it to teachers too (generalization), and a dog may learn that sitting earns treats only when a certain person gives the command (discrimination).

Shaping teaches complex behaviors by reinforcing successive approximations, steps that get closer and closer to the goal. To teach a dog to roll over, you first reward lying down, then lying on its side, then rolling partway, then the full roll. But there are limits: animals trained this way sometimes drift back toward instinctive behaviors, called instinctive drift. Raccoons trained to drop coins in a box kept rubbing the coins together instead, as they would with food.

Superstition and learned helplessness

Superstitious behavior forms when a reinforcer happens to follow a behavior that didn't actually cause it. Skinner gave pigeons food at regular times no matter what they did, and many repeated whatever they happened to be doing, like turning in circles. A player who wore certain socks during a win and keeps wearing them is showing the same thing.

Learned helplessness happens when an organism learns that it has no control over unpleasant events, so it stops trying even when escape becomes possible. Martin Seligman found that dogs exposed to shocks they couldn't escape later failed to escape shocks they easily could have. Learned helplessness has been linked to depression (5.4).

Schedules of reinforcement

Continuous reinforcement rewards every correct response. It leads to fast learning but also fast extinction once rewards stop. Partial (intermittent) reinforcement rewards only some responses; learning is slower, but the behavior resists extinction much more. Partial schedules are based on either the number of responses (ratio) or the time passed (interval), on a fixed or variable basis.

ScheduleRuleResponse patternExample
Fixed-ratioAfter a set number of responsesHigh rate, brief pause after each rewardA free coffee after every 10 purchases
Variable-ratioAfter an unpredictable number of responsesHighest, steady rate; most resistant to extinctionSlot machines; fishing
Fixed-intervalFirst response after a set amount of timeSlow after each reward, speeding up as the time nears (scalloped graph)Checking the oven near the end of baking time; studying more as a scheduled test approaches
Variable-intervalFirst response after an unpredictable amount of timeSlow, steady rateChecking for texts; pop quizzes

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1

    Negative reinforcement or punishment?

    Classify each. (a) Jada takes aspirin, and her headache goes away; she takes aspirin more often for headaches. (b) A teen loses driving privileges for texting while driving, and texts less. (c) A child gets extra chores for lying, and lies less.

    Show the solution
    1. Step 1: Ask two questions each time: did the behavior increase or decrease? Was something added or removed?
    2. Step 2: (a) Taking aspirin increases, so it's reinforcement. The headache (unpleasant) is removed, so it's negative reinforcement.
    3. Step 3: (b) Texting decreases, so it's punishment. Driving privileges (pleasant) are removed, so it's negative punishment.
    4. Step 4: (c) Lying decreases, so it's punishment. Chores (unpleasant) are added, so it's positive punishment.

    Answer: (a) Negative reinforcement, (b) negative punishment, (c) positive punishment.

  2. Example 2

    Reading a schedule from data

    Two groups of rats press a lever for food. Group A is rewarded every 5th press; Group B is rewarded after an unpredictable number of presses averaging 5. When food stops completely, Group A stops pressing after about 40 presses, and Group B keeps pressing for about 160. Name the schedules and explain the difference.

    Show the solution
    1. Step 1: Group A is rewarded after a set number of responses: fixed-ratio. Group B is rewarded after a varying number: variable-ratio.
    2. Step 2: Compare persistence: 160 ÷ 40 = 4, so Group B made four times as many presses during extinction.
    3. Step 3: With a variable schedule, the rats never know when the next reward is coming, so a long run without food doesn't signal that rewards have stopped.
    4. Step 4: This is why variable-ratio schedules are the most resistant to extinction.

    Answer: Group A: fixed-ratio; Group B: variable-ratio. The unpredictable variable-ratio schedule made the behavior about four times more resistant to extinction.

Common mistakes

  • Treating negative reinforcement as punishment. Negative reinforcement removes something unpleasant and increases behavior; punishment always decreases behavior.
  • Reading 'positive' as 'good.' Positive means something is added, whether pleasant (reward) or unpleasant (punishment).
  • Mixing up ratio and interval schedules. Ratio is about the number of responses; interval is about the time that has passed.
  • Thinking continuous reinforcement creates the most lasting behavior. It's fast to learn but fast to extinguish; variable-ratio lasts longest.

On the exam

  • Expect many scenarios that ask you to classify a consequence or a schedule, so always ask: increase or decrease? Added or removed? Number or time? Fixed or variable?
  • Graph questions may describe cumulative response lines; know that fixed-interval produces a scalloped pattern and variable-ratio a steep, steady slope.

Connected topics

Videos

  • Operant Conditioning & Reinforcement Schedules (AP Psychology Review Unit 3 Topic 8)

    Mr. SinnWatch on YouTube (opens in a new tab)

  • Unit 3B Learning Part 3 Consequences that Shape Behavior - Operant Conditioning (Updated 2026)

    Mrs. McCraryWatch on YouTube (opens in a new tab)

  • Operant conditioning: Positive-and-negative reinforcement and punishment | MCAT | Khan Academy

    khanacademymedicineWatch on YouTube (opens in a new tab)

  • Operant Conditioning in Under 3 mins (AP Psychology Unit 3 Topic 8) 3.8

    Maximum InsightWatch on YouTube (opens in a new tab)

  • Operant conditioning: Schedules of reinforcement | Behavior | MCAT | Khan Academy

    khanacademymedicineWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 3.8 Operant Conditioning. Pick an answer to see if you got it, and why.

Reinforcement scheduleMean presses per minute during trainingMean presses during a 1-hour session with no reinforcement
Fixed-ratio (every 10th press)52310
Variable-ratio (an average of every 10th press)61920
Fixed-interval (first press after each 1 minute)14180
Variable-interval (first press after an average of 1 minute)22540

Hypothetical data from rats trained to press a lever for food

Question 1 of 4

Which schedule is most similar to the payout pattern of a slot machine, and what do the data show about it?

Question 2 of 4

Which conclusion is best supported by the data?

Question 3 of 4

When 16-year-old Eli comes home an hour after curfew, his parents take away his phone for the weekend, and he starts coming home on time. Taking away the phone is an example of

Question 4 of 4

A car beeps loudly until the driver fastens her seat belt. Over time, the driver starts buckling up as soon as she gets in. The driver's behavior is being shaped by

0 of 4 answered