1. The Experiment and the Operant Model
In the 1930s, Burrhus Frederic Skinner revolutionized experimental psychology by developing the Operant Conditioning Chamber (popularly known as the Skinner Box). Unlike Ivan Pavlov's passive conditioned reflex, operant behavior is that which operates upon the environment, being selected and maintained by its consequences.
"Men act upon the world, and change it, and are changed in turn by the consequences of their actions." — B. F. Skinner
2. The Mathematics of Learning: Rescorla-Wagner
To simulate the acquisition of the association between the mechanical click of the dispenser and the delivery of food, the virtual rat's brain employs the mathematical model of Rescorla-Wagner (1972):
- ΔVi: Change in the associative strength of conditioned stimulus i (the feeder click).
- αi, β: Perceptual salience parameters and learning rate.
- λ: Maximum asymptotic value supported by the US (λ = 1.0 in the presence of food; λ = 0 in its absence).
- ΣV: Total cumulative associative expectation currently held by the organism.
3. How to Train the Virtual Rat: Step by Step
Magazine Training
The rat must learn that the metallic sound signals food. ClickDispense Pellet Manually when the rat is away from the hopper. Observe how the Sound → Food Association gauge grows. When it reaches ~60%, the rat will quickly dash to eat upon hearing the click.
Response Shaping
Deliver food when the rat approaches the left side of the chamber. Next, reinforce only when it rears up in the direction of the lever. Soon, operant probability will rise until the rat presses the bar on its own!
Intermittent Reinforcement Schedules
Once the rat is conditioned under continuous reinforcement (CRF), select more complex schedules such as Fixed Ratio (FR),Variable Ratio (VR), or Fixed Interval (FI)to observe authentic behavioral patterns recorded on the cumulative drum.
Extinction of Behavior
By shutting off the food supply in Extinction mode, observe the extinction burst: the rat will press the lever compulsively seeking reward before gradually giving up and returning to baseline exploratory levels.
4. The 5 Major Reinforcement Schedules
| Schedule | Reinforcement Rule | Cumulative Recorder Pattern | Real-World Analogy |
|---|---|---|---|
| CRF (FR-1) | Every valid response is reinforced. | Continuous, steep curve of rapid acquisition. | Turning on a light by flipping the wall switch. |
| FR (Fixed Ratio) | Reinforcement after a fixed number N of responses. | High response rate with typical Post-Reinforcement Pauses (PRP). | Piece-rate factory work or sales commission every 10 deals. |
| VR (Variable Ratio) | Reinforcement after an average variable number of N responses. | Extremely high, steady rate without pauses (steep straight line). | Slot machines and social media notification feeds. |
| FI (Fixed Interval) | First response after a fixed time t seconds. | Scallop curve: initial pause followed by terminal acceleration. | Cramming for an exam only on the night before the weekly test. |
| VI (Variable Interval) | First response after variable intervals of time. | Moderate, steady rate highly resistant to extinction. | Checking email inbox or messaging apps without a fixed schedule. |
| DRL (Low Rates) | Reinforcement delivered only if inter-response time (IRT) > t. | Very low, spaced rate (training impulse control and inhibition). | Waiting for your turn to speak in conversation without interrupting. |
5. Skinner's Cumulative Recorder
Before computers and digital monitors, B. F. Skinner invented the Cumulative Recorder: a continuous mechanical paper drum driven by an electric motor, with a pen that stepped laterally with each lever press.
The primary advantage of this device is that the animal's instantaneous response rate is directly proportional to the angular slope of the traced line:
- Flat horizontal line: Zero responses (pause or extinction).
- Gentle slope: Low rate of responding.
- Near-vertical slope: High frequency of responses per minute.
- Diagonal tick marks ( \ ): Delivery of food reward (reinforcement).
6. Stimulus Discrimination and Generalization Gradients
In operant discrimination training, an antecedent stimulus signals when a response will be reinforced (Discriminative Stimulus or SD, e.g., Light ON) versus when it is on extinction (SΔ, e.g., Light OFF). Over time, the organism learns to emit responses almost exclusively in the presence of SD.
When presenting test stimuli across a gradient of frequencies or intensities (e.g., tones from 440 Hz to 1320 Hz around the 880 Hz training tone), one obtains the Generalization Gradient: a bell-shaped curve showing how responding progressively declines as the stimulus deviates from the original SD.
7. Conditioned Emotional Response (CER) & Suppression Ratio
Developed by Estes and Skinner (1941), the Conditioned Emotional Response (CER / Conditioned Suppression) procedure demonstrates how Pavlovian fear modulates ongoing operant behavior. When a warning tone precedes an aversive grid shock, the animal acquires conditioned fear and displays freezing, temporarily suppressing lever pressing.
The magnitude of fear is measured via the Suppression Ratio (SR):
- B: Number of lever presses during the warning CS presentation.
- A: Number of lever presses during the pre-CS baseline period.
- SR = 0.50: No suppression (normal behavior, absence of fear).
- SR = 0.00: Total suppression (fear completely halted behavior).
8. The Rescorla-Wagner Model and Compound Conditioning
The Rescorla-Wagner (1972) model is the most influential formal mathematical formulation of Pavlovian associative learning. Its core tenet is that learning occurs when events in the environment violate the organism's expectations (prediction error):
- ΔVA: Associative strength update for stimulus A on this trial.
- αA: Salience / perceptibility of stimulus A (0.0 to 1.0).
- β: Learning rate determined by the unconditioned stimulus US (food or shock).
- λ: Asymptote supported by the US (λ = 1.0 with US; λ = 0.0 without US).
- ΣV: Sum of associative expectations of all stimuli present on the trial (ΣV = VTone + VLight).
Through compound summation (ΣV), the simulation reproduces canonical Pavlovian phenomena:
- Kamin's Blocking: If the Tone already fully predicts shock (VTone = 1.0), presenting [Tone + Light] followed by shock yields zero prediction error (λ - ΣV = 1.0 - 1.0 = 0). The Light fails to gain associative strength (ΔVLight = 0), beingblocked by prior learning.
- Overshadowing: When two novel stimuli are presented together, the more salient stimulus (α higher) acquires the majority of associative value, overshadowing the weaker one.
- Over-expectation: If Tone and Light are conditioned separately to V = 1.0 each, presenting both together with 1 shock causes the animal to expect 2 shocks (ΣV = 2.0). Negative prediction error causes loss of associative strength for both.
- Conditioned Inhibition (CS- / Safety Signal):When Tone predicts shock but [Tone + Light] predicts no shock, the Light develops negative associative strength (V < 0), becoming a safety signal that suppresses fear.
- Latent Inhibition: Repeated pre-exposure to Tone alone slows subsequent conditioning when the Tone is paired with shock.
9. Operant Dynamics: Spontaneous Recovery and Secondary Reinforcement
In behavior analysis, extinction does not erase earlier learning, but superimposes a fragile inhibitory memory:
- Spontaneous Recovery: After an extinction session where lever pressing ceased, a 24-hour rest period dissipates temporary inhibition, causing responding to re-emerge at the start of the next session without new reinforcement.
- Conditioned (Secondary) Reinforcers: The dispenser click sound acquires secondary reinforcing value through pairing with food. If the click sound remains enabled during extinction, responding persists longer; muting the dispenser accelerates extinction.
- Operant Punishment: When an operant response produces an aversive shock, response rate is rapidly suppressed, demonstrating the asymmetry between positive reinforcement and aversive contingencies.
10. Two-Bar Operant Chamber, Concurrent Schedules & Herrnstein's Matching Law
To study choice and time allocation, the simulation includes the Two-Bar Operant Chamber. The organism has simultaneous access to two levers (Left Bar and Right Bar), each operating under an independent reinforcement schedule (Concurrent Schedules, e.g., Conc VI 30s VI 60s).
Richard Herrnstein (1961) formulated the celebrated Matching Law, stating that the relative proportion of responses emitted on an alternative matches the relative rate of reinforcement obtained from that alternative:
- BL & BR: Number of responses emitted on the Left and Right bars.
- RL & RR: Number of reinforcements (pellets) obtained on Left and Right bars.
- Changeover Delay (COD): To prevent superstitious rapid alternation between levers, a 1.5s–3.0s penalty delay is enforced after switching levers before responses can be reinforced, ensuring genuine matching behavior.
11. Advanced Phenomena in Discrimination & Schedules
A. Peak Shift Effect (Kenneth Spence)
In intradimensional discrimination training, the organism is reinforced in the presence of one stimulus (SD = 880 Hz) and extinguished under an adjacent stimulus (SΔ = 660 Hz).
Kenneth Spence (1937) demonstrated that the resulting response gradient is thenet subtraction of the inhibitory gradient centered at SΔfrom the excitatory gradient centered at SD:
- Shift Away from SΔ: Because inhibition at 660 Hz suppresses nearby frequencies, the peak response frequency shifts away from SΔ to frequencies higher than 880 Hz (e.g., 990 Hz or 1100 Hz).
- Theoretical Significance: Empirically validates that generalization and discrimination operate as interacting vector force fields in the nervous system.
B. Ratio Strain and Operant Fatigue
When the response requirement in a Fixed Ratio schedule is increased too abruptly (e.g., jumping from CRF directly to FR-30 or FR-50), operant behavior collapses into Ratio Strain.
The organism exhibits prolonged post-reinforcement pauses, frustration-like reactions, and displacement into competing non-operant behaviors (grooming, corner sniffing). To establish high-ratio performance safely, gradual step-up is required:
12. Contextual Conditioning & The ABA Context Renewal Effect
Associative learning is deeply integrated with the environmental background (Context). The ABA Renewal paradigm is one of modern psychology's most important experimental models for clinical relapse:
Context A
Tone is paired with Shock in Chamber A. Conditioned fear develops to both the CS tone and the contextual cues of Chamber A.
Context B
The rat is moved to Chamber B and Tone is presented repeatedly without shock. Fear is successfully extinguished (SR → 0.50).
Return to Context A
Returning to Chamber A, Tone presentation triggers instant fear renewal (SR ≤ 0.35), despite no new shocks delivered.
Clinical Implications: Extinction learning (such as exposure therapy for phobias, PTSD, or substance use) is context-specific. Returning to original trauma contexts can reactivate extinguished responses.
13. Suggested Experimental Lab Protocols
Below are 4 structured protocols for experimental behavior analysis coursework:
Feeder Training & Operant Shaping
Objective: Establish autonomous lever pressing via successive approximations.
Steps: Load the Naive Rat profile. In the Shaping & Operant tab, dispense food manually until Readiness ≥ 80%. Then reinforce successive approximations until the first autonomous press.
Metric: Note total manual reinforcements and elapsed time to autonomy.
Cumulative Record Morphology: FR vs. VR vs. FI
Objective: Contrast cumulative recorder slopes across ratio and interval schedules.
Steps: Train the rat on FR-10 for 3 minutes (staircase pattern). Switch to VR-10 (smooth steep line). Switch to FI-20s (scallop acceleration).
Metric: Calculate responses per minute in each schedule.
Verification of Herrnstein's Matching Law
Objective: Empirically test choice allocation on concurrent schedules.
Steps: Enable the Two-Bar Chamber with Conc VI 20s VI 40s. Run for 5 minutes at 10x simulation speed.
Metric: Compare response ratio BL / (BL + BR) with reinforcement ratio RL / (RL + RR).
Kamin's Blocking Effect (Pavlovian Blocking)
Objective: Demonstrate that prediction error governs associative learning.
Steps: In the Classical Conditioning tab, run the Kamin Blocking protocol.
Metric: Verify trial-by-trial associative strength and confirm why VLight remains near zero.
14. Canonical Glossary of Experimental Analysis of Behavior (EAB)
| Term / Acronym | Full Name | Conceptual Definition | Manifestation in Simulation |
|---|---|---|---|
| SD | Discriminative Stimulus | Antecedent stimulus in whose presence a response is reinforced. | Light or Tone ON signaling that lever press yields food. |
| SΔ | S-Delta | Antecedent stimulus in whose presence a response is on extinction. | Absence of light/tone; lever pressing produces no consequence. |
| US | Unconditioned Stimulus | Biologically potent stimulus eliciting an innate reflex. | Grid shock (pain) or food pellet (primary reward). |
| CS | Conditioned Stimulus | Initially neutral stimulus that acquires function through pairing with US. | 880 Hz tone eliciting freezing after shock association. |
| UR / CR | Unconditioned / Conditioned Response | Innate reflex (UR) vs. learned response to CS (CR). | Innate startle to shock (UR) vs. anticipatory freezing to tone (CR). |
| SR | Suppression Ratio | Estes-Skinner index for objective measurement of conditioned fear. | 0.50 (normal/no fear) down to 0.00 (complete suppression). |
| PRP | Post-Reinforcement Pause | Transient cessation of responding immediately following reward. | Horizontal flat pause on cumulative drum in FR and FI schedules. |
| IRT | Inter-Response Time | Elapsed time interval between two consecutive responses. | Monitored in DRL schedule to enforce slow, paced responding. |
| COD | Changeover Delay | Penalty delay enforced upon switching levers in concurrent schedules. | Prevents accidental reinforcement of rapid switching behavior. |
| Scallop | Scallop Curve | Accelerating rate of responding across a fixed interval. | Curvilinear pattern characteristic of Fixed Interval (FI) records. |
| Extinction Burst | Extinction Burst | Abrupt temporary rise in response rate when reinforcement is halted. | Rapid compulsive bar pressing immediately after food cut-off. |
| Spontaneous Recovery | Spontaneous Recovery | Partial return of an extinguished response after a rest interval. | Return of lever pressing after 24h rest with no new reward. |