Elderly Priming: Did It Ever Work?

The Short Version

Bargh, Chen, and Burrows (1996) reported that priming students with words related to old age made them walk more slowly, with effect sizes above one standard deviation (d = 1.04 and .79) across two studies of just 30 participants each. The finding became a cornerstone of behavioral-priming research and a famous casualty of the replication crisis after Doyen et al. (2012) failed to reproduce it. The Doyne et al. (2012) article came at the right time to cause a paradigm shift, but it was not the first replication failure.

The first replication failure appeared only a couple of years after the 1996 article: Dijksterhuis et al. (1998) found the same comparison at roughly a quarter of a standard deviation (d = .25 and .29), nonsignificant, and reinterpreted the disappearance as a side condition for a new contrast effect. Because social psychologists tracked the sign of an effect rather than its magnitude, a series of studies that did not reproduce Bargh’s result were published as successful demonstrations of moderators instead of as replication failures.

By 2012, elderly priming had been reported in 8 articles with 12 studies and 15 tests. I show that a meta-analysis at this time would have shown a wide prediction interval with possible effect sizes ranging from -0.1 to +1.0. Thus, Doyen et al.’s replication failure was entirely consistent with the full existing evidence. It was therefore entirely reasonable for Kahneman (2012) to ask for new and stronger evidence that Bargh never delivered. Since then, independent preregistered studies large failed to replicate past effects and priming theorists have largely abandoned priming research. The death of the priming paradigm provides a valuable lessons about the need to build theories on robust empirical foundations to avoid investing resources on phenomena that do not exist.

The Long Version

John A. Bargh studied with Robert Zajonc at the University of Michigan, earning his Ph.D. in 1981. Zajonc was an early champion of unconscious processes and used masked, subliminal presentations to study influences occurring outside conscious awareness. At New York University, Bargh developed a related program on automatic social cognition. At the time, studies from several laboratories suggested that stimuli presented outside awareness could influence feelings, judgments, and immediate reactions to other stimuli. Some of this evidence has since been challenged: meta-analyses find many subliminal effects hard to replicate, and concerns have been raised about publication bias and about whether participants were ever fully unaware of the masked stimuli.

Bargh matters for the history of psychology because his 1996 article made a much stronger claim. Rather than showing that unnoticed stimuli can shift immediate perceptions or evaluations, Bargh and colleagues claimed that activating a social concept could have a lasting influence on overt behavior without people being aware of that influence. In their famous elderly-priming experiment, students completed a scrambled-sentence task in which, in one condition, several words were related to old age and, in the control condition, were not. The words were visible, but participants were presumably unaware of the manipulation’s purpose or its possible effect on their behavior. Told to go to another room for a second study, participants then had their walking speed down the hallway measured as the real dependent variable. Bargh et al. reported that a few old-age words made students walk more slowly than controls.

The difference was not small. Effect sizes for two-group differences are often expressed in standard-deviation units. Cohen (1988) classified d = .50 as a medium effect and d = .80 as large. For comparison, an IQ test has a standard deviation of about 15 points, so half a standard deviation is 7.5 IQ points. In Bargh et al.’s first study, Experiment 2a, the effect was slightly larger than a full standard deviation, d = 1.04 — an exceptionally large effect for such a subtle manipulation.

Psychologists rarely run direct replications, but Bargh et al.’s article contained one. Experiment 2b used essentially the same procedure; again, primed participants walked significantly more slowly, and the effect stayed large, d = .79. The article thus appeared to offer unusually convincing evidence: two independent studies, same procedure, both producing significant and large effects.


The 1996 article spawned a large literature using primes such as money or God to influence behavior. Then, in 2012, Doyen et al. reported that they could not replicate the elderly-priming effect and suggested the original findings might be, at least partly, methodological artifacts. Walking speed in the original studies was timed manually with a stopwatch. And although the person timing walking speed was blind to condition, Doyen et al. noted it was unclear whether the experimenter who administered the priming task was also blind. If experimenters knew which participants had been primed, their expectations could have shaped participants’ behavior.

Doyen et al. tested this in a second experiment by manipulating experimenters’ expectations directly: some were led to expect primed participants to walk more slowly, others to walk faster. Objectively measured walking speed tracked those expectations — the elderly-prime effect appeared only when experimenters expected slowing — and the manual stopwatch measurements tracked them even more strongly. Doyen et al. thus provided experimental evidence that researchers’ expectations could produce the outcome of a behavioral-priming study.

Bargh responded in March 2012 with a Psychology Today post titled “Nothing in Their Heads.” He opened by noting that the 1996 finding had been theoretically predicted and fit a growing literature on automatic influences on behavior, but much of the post attacked the quality of Doyen et al.’s work and the peer review at PLOS ONE. Knowledgeable social-psychology editors and reviewers, he argued, would have caught the methodological problems; he had not been asked to review the paper, and implied that had he done so, he would have recommended rejection.

The tone drew wide criticism (Comments). Srivastava (2012), for one, characterized the episode as Bargh going “bananas.” The attack on the journal prompted a reply from the PLOS ONE editors on Bargh’s blog post (cf. Hodgkinson, 2012; Srivastava, 2012). Bargh later expressed regret over the tone and took the post down (Bartlett, 2013).

Bargh’s Defense of His Results in 2012

Bargh’s first substantive defense concerned experimenter effects. He stated that the experimenter in the original study had been blind to hypothesis and condition, and that a different person, also blind to condition, measured walking speed. If accurate, this substantially weakens Doyen et al.’s suggestion that experimenter expectations produced the original findings. But it does nothing for the more basic result: Doyen et al. used a larger sample and objective measurement and still failed to reproduce the effect.

Bargh therefore proposed several procedural differences to explain the replication failure. First, he argued that Doyen et al. had drawn participants’ attention to walking by telling them to “go straight down the hall when leaving,” and that making an automatic behavior conscious could eliminate the priming effect. But Doyen et al.’s article contains no such instruction; it says only that participants were “clearly directed to the end of the corridor” — and Bargh et al.’s own procedure had likewise directed participants toward the elevator down the hall. It is unclear the difference existed at all.

Second, Bargh questioned the strength of the manipulation: Doyen et al. put an elderly-related word in all 30 scrambled-sentence items, and Bargh argued that too many related words could make participants consciously aware of the theme and cancel an unconscious effect. Doyen et al. did find some evidence that participants could identify the elderly theme when directly probed — but nearly all denied noticing any connection between the sentence task and their walking.

Third, Bargh appealed to culture: priming works only if the association already exists in participants’ minds, so Belgian students might not link old age with slowness as American students do. Possible in principle — but Doyen et al. had adapted their materials to the Belgian sample by surveying 80 people about concepts associated with old age and selecting the frequent responses. Bargh et al.’s original article, by contrast, never measured whether their own participants associated the elderly stereotype with slower walking.

Finally, Bargh appealed to the accumulated literature: stereotype and behavioral priming had been replicated many times, so it was unreasonable to doubt the phenomenon over a single failure. This is persuasive only if the published record is an unbiased sample of all experiments run. That was precisely what the emerging replication crisis called into question. Journals favored novel, significant results and rarely published failed replications, so the sheer number of published successes could not reveal how often behavioral-priming experiments actually worked.

Doyen (2012) Was Not the First Failure

Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it, Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it,

It is a bit of a shame that many of the arguments Bargh uses in his criticism of the Doyen study are arbitrary, unsupported and, on occasion, false in light of other research (even some from the area of priming) (Nordbeck, 2012).

The earlier failures had been published, but not as failures. They appeared as successful demonstrations of moderators: conditions under which the effect was supposed to grow, shrink, or reverse (Cesario et al., 2006; Dijksterhuis et al., 1998; Hull et al., 2002). Bargh later cited these very studies without noting that none of them had reproduced his basic result — a difference between an elderly prime and a control group.

That this could happen reflects a habit of the field. Social psychologists tracked patterns of statistical significance more than the magnitude of effects. Evidence for a moderator requires a significant interaction, and once an interaction turns up, attention shifts from the main effect to the conditional effects that explain it. The claim is no longer “priming works” but “priming works differently under condition X” — and a study can support that claim while quietly failing to reproduce the main effect it was built on.

Dijksterhuis et al. (1998) is the clearest case, appearing just two years after the original. Studies 2a and 2b each contained the two conditions needed to test elderly priming — an elderly prime and a neutral prime, followed by the same judgment task. In both, primed participants walked slightly more slowly, but the effects were small and nonsignificant, d = .25 and .29, against Bargh’s d = 1.04 and .79. Same direction, a quarter of the magnitude.

But testing Bargh was not the point of their paper; behavioral contrast was. Some primed participants also judged a specific elderly exemplar — the Dutch Queen Mother, Princess Juliana — and then walked faster than both the neutral and the ordinary elderly-prime groups. That contrast effect was significant in both studies and became the headline. The authors did notice that their ordinary elderly prime had failed to reproduce Bargh’s assimilation effect, called it “somewhat surprising,” and proposed that the intervening judgment task had wiped it out — reading the small same-direction means as “residual assimilation.” The failure was not overlooked; it was reinterpreted into a footnote. By this reading, the first evidence that elderly priming is not as large or robust as advertised appeared fourteen years before Doyen.

Their explanation is itself revealing. If inserting an innocuous judgment task can cut an effect by roughly 75%, then the effect is extraordinarily sensitive to procedural detail. Researchers wanted theoretically predicted moderators that would specify when and why priming occurs. What the evidence kept pointing to instead were unknown moderators: incidental features of a procedure that swing the effect from large to nothing. A phenomenon that surfaces only under a narrow and poorly understood set of conditions is not the robust effect Bargh’s original experiments described.

Dijksterhuis was only the first. Hull et al. (2002) ran the paradigm in two studies and found priming only among participants high in self-consciousness; the overall effects were moderate but nonsignificant, ds = .57 and .54, in small samples. Cesario et al. (2006) added a youth prime that sped participants up, but the elderly-versus-control difference was small and nonsignificant, d = .23. Jeffries and Fazio (2008) found an interaction with a stopping rule on an anagram task and no main effect at all, d = −.05. N. Wyer (2011) produced the first robust replication — d = .70 on walking speed and d = 1.03 on working memory. And in 2012, the same year as Doyen, one article reported strong effects in one study (ds = 1.26, 1.05) but weak ones in another (ds = .29, .31).

Seen in this company, Doyen’s failure is unremarkable. Some studies produce large estimates and others produce weak estimates and all estimates in small samples have large sampling error and make estimation of the true effect size impossible. The common solution to imprecise estimates in small studies is to combine the studies in a meta-analysis.

A Meta-Analysis of Elderly Priming

Existing meta-analyses of priming pool many kinds of primes and outcomes, and the resulting heterogeneity tells us little about any one paradigm. So I conducted a meta-analysis restricted to elderly priming: 15 tests from 12 studies in 8 articles. Because publication bias is evident in the broader literature, I used a selection model that estimates the bias and returns a corrected effect size (Vevea & Hedges, 1995), with clustered bootstrapping to handle multiple tests nested within articles and to build the confidence and prediction intervals.

The bias-corrected effect was d = .42, 95% CI [−.01, .71]. The point estimate is close to the much larger meta-analysis of Dai et al. (2023), but the interval is wide and includes zero. The studies are small and individually imprecise, and the heterogeneity is itself barely pinned down: tau = .26, 95% CI [.00, .37]. Nothing here compels the conclusion that Bargh’s and Doyen’s results reflect different true effects.

The estimated mean and variation of the population effect sizes produce a 95% prediction interval that ranges form d = -0.10 to d = 1.03.

Thus, a replication study so large that sampling error is negligible could land anywhere from essentially nothing to a very large effect, because the existing evidence simply does not fix the size of the effect.

The meta-analysis also cannot explain how Bargh got two significant results from samples of N = 30. If the true effect were d = .42, a study that small has only about a 20% chance of reaching significance, so two independent significant results would occur about 4% of the time. Either Bargh was lucky, or something in his particular setup inflated the effect — for instance, a confound in the word list, if “Florida” slows NYU students’ walking by evoking a laid-back-Southern stereotype rather than an elderly one.

In short, a meta-analysis of the studies available in 2012 would have shown Bargh that Doyen’s result was entirely consistent with the broader evidence, however much it clashed with his own. There was no need to reach for incompetence. His deeper error was assuming that his two original studies, together with conceptual replications using other primes, had already established a general effect of elderly priming on behavior.

An important caveat is that publication-bias tests and corrections are difficult, especially when the set of studies is small. After Doyen et al.’s article was published, researchers shared anecdotal reports of additional unpublished replication failures (OSF Google Groups). If these unpublished studies could be recovered and included, they would likely reduce the average effect-size estimate.

Source: Comments to Ed Young’s Article
Source: Comments to Ed Young’s Article

The Fall-Out

Concerns about the credibility of social psychology were already building. Daniel Kahneman believed in priming and had featured it prominently in Thinking, Fast and Slow, yet he was growing worried about the field’s reputation and raised it with Bargh directly (Bartlett, 2012). In a widely circulated email to Bargh and other leading priming researchers, he proposed a simple remedy: if Doyen et al.’s failure was a fluke or a botch, the experts could settle the matter by running rigorous replications and showing the effects held. With an expected effect of d = .40, pooling resources for N = 400 would give roughly 98% power to detect it.

The collaboration never happened. Independent researchers ran preregistered replications instead, often with disappointing results — in Dai et al.’s (2023) meta-analysis, the preregistered studies, most of them replications, produced essentially no effect at all (d = .02).

In 2017, Kahneman returned to the problem, citing Overall (1969) on why underpowered research is not merely wasteful but pernicious: it inflates the share of false positives among published findings. Bargh’s experiments, with 15 participants per condition, had little power to detect any plausible effect and cannot account for the near-90% success rates in the published priming literature. Hundreds of studies manufactured the appearance of a robust phenomenon that low power, publication bias, heterogeneity, and near-zero preregistered effects do not support.

So the classic behavioral-priming paradigms simply faded from the research frontier. In that sense the field is practically dead — a stark case of strong belief outrunning weak evidence. Researchers can always find a reason to dismiss a failure and build ever more elaborate theory around inconsistent results, all without first establishing that the phenomenon the theory explains is real. As the meta-analysis above shows, even the 2012 evidence carried enough uncertainty to justify Kahneman’s call.

Which raises the obvious question: why wouldn’t a true believer, with the resources of a Yale professor, simply run the study and prove the point? Bartlett (2013) put it to Bargh directly:

So why not do an actual examination? Set up the same experiments again, with additional safeguards. It wouldn’t be terribly costly. No need for a grant to get undergraduates to unscramble sentences and stroll down a hallway.

Bargh was unenthusiastic. He wouldn’t ask graduate students worried about their job prospects to spend time on stigmatized research, and he was aware that some critics thought he had a “special touch” with priming — a word that sounds like praise but isn’t. “I don’t think anyone would believe me,” he said.

Source: Comments on Ed Young Article

The same article also reports an interview with Axel Cleeremans, one of Doyen’s collaborators. Cleeremans reports that his collaborators and he found found the tone of Bargh’s letter insulting, but that they also took his criticism seriously enough to conduct another replication study that addressed his stated concerns. The study also failed to replicate the original results. Another unpublished replication failure was posted online by Pashler et al. (2011; Srivastava, 2012)

While the evidence that elderly priming works was mixed and inconclusive, Bargh expressed strong confidence in his findings and his theory of unconscious activation of behavior. In his blog post, he believed that the long-term future of priming research was secure because the phenomenon rested on a large and converging empirical foundation:

Like many scientists, I take the long view and continue to have faith in science as a cumulative process, with particular studies that vary in their methods and approaches converging on deeper underlying principles and mechanisms. No single experiment, standing by itself, should be the basis for concluding anything, maybe most especially in psychological research. And when a single study does not replicate another one whose findings are solidly embedded in theories of more than one scientific field and which is consistent with dozens if not hundreds of other conceptual replications, then responsible scientists—and responsible science journalists—do not rush to judgment and make claims that the entire phenomenon in question is illusory. Thomas Kuhn famously argued in The Structure of Scientific Revolutions that it is the accumulation of contrary evidence over time, not single studies, which is needed to overturn established concepts and principles in any branch of science.

A decade on, his faith in cumulative science has been vindicated — just not as he expected. Failure after failure, and no convincing new evidence from the original researchers, has left claims about unconscious behavioral priming no better established than Freud’s speculations about the unconscious. Priming researchers published and circulated their successes and buried their failures, manufacturing the look of robustness; and as shown here, even the successes never told a clean story or gave clear evidence that elderly priming had ever produced a reliable effect.

Bargh was right about one thing: no single failure should have settled it. What mattered was the accumulation of evidence. That accumulation just happened to vindicate Doyen’s skepticism rather than his own confidence. Maybe, somewhere in his unconscious, Bargh knows this too — which might be why he never tried to replicate his 1996 findings.

The Lesson

Bargh is not the first scientist to spend much of a career on an idea that never lived up to its promise. That is an occupational hazard of working at the edge of a field: the questions are hard, the methods imperfect, and a genuine breakthrough is difficult to tell from a false lead until many technical problems are solved. Strong belief supplies the motivation to push through those obstacles — but it also makes the warning signs of failed studies easier to miss. Economists call inflated asset prices “irrational exuberance”: a stock can stay overvalued for as long as enough investors share the illusion, then crash to its true worth. Psychological theories kept aloft by shared belief rather than evidence behave the same way, as the citation history of Bargh’s 1996 article shows.

That casual attitude toward evidence is on display in Matthew Lieberman’s 2012 blog post on Doyen’s failure — a post that mentions, in passing, an unpublished non-replication from Lieberman’s own lab:

There have been multiple unpublished non-replications of the elderly-walking (including one in my lab that I may discuss in a future blog), but it is hard to know what to make of them given that they haven’t been peer-reviewed.

The catch-22 is clean. We trust only peer-reviewed findings, but psychology journals publish overwhelmingly significant results (Sterling, 1959; Sterling et al., 1995). So the failures stay invisible — unless they are dressed up as significant interactions, or land in a venue like PLOS ONE that isn’t gatekept by the same community. More telling still is what Lieberman took the stakes to be:

At a certain level it does not matter whether the exact primes Bargh used produce a change in walking speed over the exact distance he measured. Some have said ‘We need to replicate this exactly. Conceptual replications aren’t good enough’. But I’m not sure why we care about this specific manipulation unless we are about to start using it as an intervention to treat patients. What we care about is whether priming-induced automatic behavior in general is a real phenomenon. Does priming a concept verbally cause us to act as if we embody the concept within ourselves? The answer to this question is a resounding yes.

The “resounding yes” rests entirely on a published record that reports successes and buries failures — and, worse, treats interaction effects that masked the very replication failures at issue as further confirmation. Shown only the hits, it is easy to conclude the phenomenon must be real. Kahneman (2017) later admitted he had trusted that record too much when he built on it in his book. The lasting contribution of the controversy is that younger psychologists now treat direct replication as something more than an optional courtesy.

Younger scientists can lower the risk by studying the previous generation’s mistakes. We are fond of saying that each generation stands on the shoulders of giants. We say less often that many equally gifted scientists climbed just as high and backed ideas that proved wrong. The ones whose theories survived were not necessarily wiser or more careful at the outset; to a real extent, nature happened to cooperate. Darwin’s mechanism of natural selection withstood ever more stringent tests; Lamarck’s mechanism of inheritance did not. Neither man could have known which way it would go when he set out.

Because you cannot know in advance whether nature will cooperate, the main lesson of what Kahneman (2012) called the “train wreck” of priming research is methodological before it is anything else: do not expect clean answers from small samples, and do not trust significant results from journals that never publish the null ones. And beyond method, a matter of temperament — stay skeptical, and resist the pull to become a believer. As Feynman put it:

The first principle is that you must not fool yourself, and you are the easiest person to fool.

1 thought on “Elderly Priming: Did It Ever Work?

Leave a Reply