Famous psychology experiments — and how well each one held up
The best-known psychology experiments vary enormously in quality. Pavlov, Asch, Loftus and Ebbinghaus have held up. Milgram, the Bobo doll, the bystander effect and the marshmallow test are real but routinely overstated. The Stanford Prison Experiment, Little Albert and Rosenhan should not be cited as evidence at all. Here is each one, rated.
Psychology’s famous experiments are famous for different reasons. Some are load-bearing foundations. Some are cautionary tales that survive in textbooks long past their evidence. A list that treats them all the same is worse than no list — so every entry here carries an honest rating.
The ones that held up
Pavlov’s dogs (1890s–1900s)
What it did: Ivan Pavlov paired a neutral signal with food until the signal alone produced salivation — classical conditioning.
How it held up: Foundational and thoroughly replicated. It underpins how a great deal of fear is learned, and why exposure therapy works — though on the current inhibitory learning model the original fear association is not erased. New learning (“this cue no longer predicts danger”) develops alongside it and competes with it, which is why fear can return in new contexts.
Asch’s conformity experiments (1951–1956)
What it did: Participants judged which line matched a reference line while confederates gave obviously wrong answers. A large share went along with the group at least once.
How it held up: Robust and replicated across decades. The under-quoted finding matters more than the famous one: a single dissenting ally collapses conformity dramatically. Being the one person who says something is one of the most powerful social acts available.
Loftus and Palmer, the car crash study (1974)
What it did: Witnesses to a filmed collision were asked how fast the cars were going when they “smashed” versus “hit” each other. The verb changed their speed estimates — and made some falsely recall broken glass.
How it held up: Repeatedly replicated and consequential. It established that memory is reconstructive, and it has reshaped how police interviews and lineups are conducted.False memories →
Tversky and Kahneman on heuristics and biases (1974)
What it did: A programme of studies showing that judgement under uncertainty runs on shortcuts — availability, representativeness, anchoring — that produce predictable errors.
How it held up: Launched behavioural economics and won a Nobel prize. Individual sub-effects vary in size, but the core insight is solid.The biases guide →
Ebbinghaus and the forgetting curve (1885)
What it did: Hermann Ebbinghaus memorised nonsense syllables and tested himself over time, charting how fast forgetting happens.
How it held up: Replicated by Murre and Dros in 2015 to within a few percent — 130 years later. Forgetting is steepest in the first hours, then flattens.
The ones that are real but routinely overstated
Milgram’s obedience experiments (1961–1963)
What it did: Participants were instructed to deliver what they believed were escalating electric shocks to another person. A striking proportion continued to the maximum.
How it held up: The rates are real and have been partially replicated. The interpretation is disputed: “engaged followership” accounts argue participants complied because they were persuaded they were helping science, not because humans blindly obey. That is a meaningfully different lesson than the one usually drawn.
Bandura’s Bobo doll study (1961)
What it did: Children who watched an adult attack an inflatable doll were more likely to attack it themselves.
How it held up: Observational learning is well supported. The leap from hitting an inflatable toy designed to be hit, to real-world aggression, is not supported by this study — and it is the leap most citations make.
Darley and Latané on the bystander effect (1968)
What it did: People were slower to help when they believed others were present — diffusion of responsibility.
How it held up: The effect is real and replicated. But the story attached to it is not: the Kitty Genovese case that inspired it was substantially misreported — the claim of 38 witnesses who watched and did nothing was inaccurate, and the newspaper that published it later acknowledged serious flaws. Also, meta-analytic work finds that in genuinely dangerous situations, bystanders often do intervene.
Harlow’s rhesus monkeys (1958)
What it did: Infant monkeys separated from mothers preferred a soft cloth surrogate over a wire one that dispensed milk — comfort mattered more than feeding.
How it held up: The finding was important and shifted thinking on attachment. The studies were also ethically indefensible by any modern standard, and could not be run today.
The marshmallow test (1972; replication 2018)
What it did: Children who waited for a second treat rather than eating one immediately were later reported to have better outcomes.
How it held up: A larger, better-controlled replication found the association was much smaller once family background and socioeconomic status were accounted for. It measures circumstances at least as much as willpower — a child who has learned that promised treats do not arrive is being rational, not impulsive.
Ego depletion and the radish experiment (1998; failed replication 2016)
What it did: Resisting a temptation was said to drain a limited pool of willpower, impairing later self-control.
How it held up: A large pre-registered multi-lab replication found an effect close to zero. The concept is not proven fake, but it can no longer be stated as established. This is the clearest example of an idea that reached bestseller status ahead of its evidence.Willpower, honestly →
The ones that should not be cited as evidence
The Stanford Prison Experiment (1971)
What it did: Students assigned to be guards or prisoners in a mock prison; the study was halted early amid reports of escalating cruelty.
How it held up: Heavily criticised on method and integrity. Guards were briefed toward the desired behaviour, the sample was self-selected by an advertisement mentioning prison life, there was no control condition, and it was a demonstration rather than a controlled experiment. It does not show that situations automatically turn ordinary people cruel — and it is still taught as if it does.
Little Albert (1920)
What it did: Watson and Rayner conditioned an infant to fear a white rat by pairing it with a loud noise.
How it held up: One infant, no control, informal outcome measures, and ethics that would end a career today. Historically important as an origin story; evidentially close to worthless.
Rosenhan’s “On Being Sane in Insane Places” (1973)
What it did: Healthy pseudopatients reportedly gained psychiatric admission by claiming to hear a word, then behaved normally and were kept in anyway.
How it held up: Enormously influential on psychiatric reform — and subsequently subject to serious doubt. Later investigative work found records inconsistent with the published account and was unable to verify most of the participants. Whatever its historical effect, it should not be presented as sound evidence.
In 2015 the Open Science Collaboration re-ran 100 studies from leading journals. 97% of the originals reported significant results; 36% of the replications did. Gilbert and colleagues published a formal rebuttal in 2016 arguing the project underestimated true reproducibility — and that argument is itself worth reading. The honest summary sits between them: many findings are solid, a meaningful minority were oversold, and the field responded with pre-registration, bigger samples and open data. A science that audits itself in public is working, not failing.
How to read any famous experiment
- How many participants? One infant is an anecdote.
- Was there a control condition? Demonstrations are not experiments.
- Has anyone repeated it? Independent replication outranks fame.
- Does the conclusion travel further than the design? Hitting a doll is not violence; a mock prison is not a prison.
- Who benefits from the popular version? The catchier the takeaway, the harder to dislodge.
Frequently asked
What is the most famous psychology experiment?
Was the Stanford Prison Experiment real?
Which psychology experiments could not be done today?
What is the replication crisis?
Are famous psychology experiments still useful?
Which experiments have held up best?
Sources & further reading
- Craske, Treanor, Conway, Zbozinek & Vervliet (2014), Behaviour Research and Therapy 58, 10–23 — the inhibitory learning account of exposure: the original fear association survives, and treatment builds a competing one. Craske et al., 2014 ↗
- Open Science Collaboration (2015), Estimating the reproducibility of psychological science, Science. Science, 2015 ↗
- Gilbert, King, Pettigrew & Wilson (2016), Comment on the above — the counter-argument. Science, 2016 ↗
- Studies named on this page are attributed inline by author, year and journal. Full linked citations live on the topic pages in each wing, and every debunked claim is sourced on our myths page.