Elderly Priming: Did It Ever Work?

John A. Bargh studied with Robert Zajonc at the University of Michigan, where he received his Ph.D. in 1981. Zajonc was an early proponent of the importance of unconscious processes and used masked, subliminal presentations to study influences that occurred outside conscious awareness. After moving to New York University as an assistant professor, Bargh developed a related research program on automatic social cognition. At the time, numerous studies from several laboratories suggested that stimuli presented outside participants’ awareness could influence feelings, judgments, and immediate reactions to other stimuli.

For the history of psychology, however, Bargh is especially important because his 1996 article made a much stronger claim. Rather than showing that unnoticed stimuli can influence immediate perceptions or evaluations, Bargh and colleagues suggested that activating a social concept could have a lasting influence on overt behavior without people being aware of this influence. In their now famous elderly-priming experiment, students completed a scrambled-sentence task. In one condition, several of the words were related to old age; in the control condition, they were not. The words themselves were visible, but participants were presumably unaware of the purpose of the manipulation and of its possible influence on their behavior. After completing the task, participants were told to go to another room for a second experiment. Their walking speed down the hallway was the real dependent variable. Bargh et al. reported that merely encountering a few words related to old age made students walk more slowly than students in the control condition.

Importantly, the observed difference between the groups was not small. Effect sizes for differences between two groups are often expressed in standard deviation units. Cohen (1988), for example, classified a standardized mean difference of d = .50 as a medium effect and d = .80 as a large effect. For comparison, the standard deviation of an IQ test is about 15 points, so half a standard deviation corresponds to 7.5 IQ points. In Bargh et al.’s first elderly-priming study, Experiment 2a, the estimated effect was slightly larger than one standard deviation, d = 1.04. This was an exceptionally large effect for such a subtle manipulation.

Psychologists rarely conduct direct replications, but Bargh et al.’s article contained one. Experiment 2b used essentially the same procedure as Experiment 2a. Once again, participants exposed to elderly-related words walked significantly more slowly, and the estimated effect remained large, d = .79. Thus, the original article appeared to provide unusually convincing evidence: two independent studies using the same procedure both produced statistically significant and large effects.

The more damaging implication of Doyen et al.’s article was that Bargh et al.’s original findings might have been produced, at least partly, by methodological artifacts. In the original studies, walking speed was measured manually with a stopwatch. Moreover, although the person measuring walking speed was blind to experimental condition, Doyen et al. noted that it was unclear whether the experimenter who administered the priming task was also blind. If experimenters knew which participants had received the elderly primes, their expectations could have influenced participants’ behavior. Doyen et al. therefore suggested that experimenter expectations offered a possible alternative explanation for the original findings.

Doyen et al. tested this possibility in a second experiment by deliberately manipulating experimenters’ expectations. Some experimenters were told to expect primed participants to walk more slowly, whereas others were told to expect them to walk faster. Participants’ objectively measured walking speed differed depending on experimenter expectations: the elderly-prime effect appeared when experimenters expected participants to slow down but not when experimenters expected them to speed up. Manual stopwatch measurements were even more strongly influenced by experimenter expectations. Thus, Doyen et al. provided experimental evidence that the expectations of researchers could influence the outcome of a behavioral-priming experiment.

Bargh responded in March 2012 with a Psychology Today blog post titled “Nothing in Their Heads.” He began by emphasizing that the original 1996 finding had been predicted on theoretical grounds and was consistent with a growing literature on automatic influences on behavior. But much of the post focused on the quality of Doyen et al.’s research and the peer-review process at PLOS ONE. Bargh argued that knowledgeable editors and reviewers in social psychology would have recognized methodological problems in the replication. He noted that he had not been asked to review the paper and suggested that, had he reviewed it, he would have identified these problems before publication. Contemporary coverage characterized the post as an unusually aggressive response to a replication failure.

Bargh’s first substantive defense concerned experimenter effects. He stated that the experimenter in the original study had been blind to the hypothesis and condition and that a different person, who was also blind to condition, measured walking speed. If this description of the original procedure is accurate, it substantially weakens Doyen et al.’s suggestion that experimenter expectations produced Bargh et al.’s original findings. However, it does not explain the more basic result: Doyen et al. used a larger sample and objective measurement and failed to reproduce the original effect.

Bargh therefore proposed several procedural differences that he believed explained the replication failure. His first argument was that Doyen et al. had drawn participants’ attention to walking by instructing them to “go straight down the hall when leaving.” According to Bargh, drawing conscious attention to an otherwise automatic behavior could eliminate the priming effect. However, Doyen et al.’s article does not contain this instruction. It states only that participants were “clearly directed to the end of the corridor.” Moreover, Bargh et al.’s original procedure also directed participants toward the elevator down the hall. Thus, it is unclear that this procedural difference actually existed.

Bargh’s second argument concerned the strength of the priming manipulation. Doyen et al. included an elderly-related word in all 30 scrambled-sentence items, whereas Bargh argued that using too many related words could make participants consciously aware of the theme and thereby eliminate an unconscious priming effect. Doyen et al. did find some evidence that participants could identify the elderly theme when directly tested, although nearly all participants denied recognizing a connection between the sentence task and their subsequent walking behavior.

Bargh’s third argument concerned cultural differences. Priming can influence behavior only if the relevant association exists in participants’ minds. He therefore suggested that Belgian students might not associate old age with slowness in the same way as American students. This explanation is possible in principle, but Doyen et al. had adapted their materials to the Belgian population by surveying 80 participants about concepts associated with old age and selecting frequently mentioned responses. Bargh et al.’s original article, in contrast, did not directly measure whether their participants associated the elderly stereotype with slower walking.

Finally, Bargh appealed to the accumulated literature. He argued that stereotype priming and behavioral priming had already been replicated many times, making it unreasonable to question the phenomenon on the basis of one unsuccessful replication. This argument sounds persuasive if the published literature is treated as an unbiased record of all experiments that were conducted. But that assumption was precisely what the emerging replication crisis called into question. Journals strongly favored novel and statistically significant results, whereas unsuccessful replication attempts were difficult to publish. Consequently, the number of published successes could not reveal how often behavioral-priming experiments actually succeeded.

More importantly, Doyen et al.’s study was not the first indication that Bargh et al.’s elderly-priming effect was difficult to reproduce. Two prominent earlier studies—Hull et al. (2002) and Cesario et al. (2006)—had been presented as successful replications because they obtained theoretically interesting moderator effects. Hull et al. reported that elderly priming slowed walking mainly among participants high in private self-consciousness, whereas Cesario et al. reported that the effect depended on participants’ attitudes toward elderly people. Bargh himself later cited these studies as evidence that the elderly-priming effect had replicated.

However, this interpretation focuses on statistically significant interactions and conditional effects rather than on replication of the original main effect. Cesario et al., for example, reported only a near-significant overall effect of priming condition, F(2, 64) = 2.93, p = .06. Once the question is changed from “Can we find some condition under which elderly priming affects walking?” to the original question—“Does elderly priming make people walk more slowly?”—the history of the finding looks quite different.

I think that last sentence is the right place to stop this section. It sets up the next section, where you can introduce the main-effect/simple-effect distinction and then bring in effect size. That is where the argument becomes more interesting than the familiar Bargh–Doyen story: Doyen was not the first replication failure at all.

Overlooked Replication Failures

Social psychologists traditionally focused more on patterns of statistically significant results than on the magnitude of effect sizes. After Bargh et al. (1996) provided evidence that behavioral priming works, subsequent researchers naturally tried to build on this finding rather than merely repeat it. One way to advance a finding is to identify moderators: conditions under which an effect becomes stronger, weaker, or even reverses. Statistically, evidence for moderation requires an interaction effect. Once an interaction is found, attention typically shifts from the overall main effect to the conditional effects that explain the interaction. The theoretical claim is no longer simply that “priming works,” but that priming works differently under particular conditions.

This way of thinking helps explain how failures to reproduce Bargh et al.’s original effect could be published in leading social-psychology journals without being regarded as replication failures. The first examples appeared only two years after the original article, in Dijksterhuis et al. (1998). Their Studies 2a and 2b closely paralleled one another and included the two conditions needed to test the original elderly-priming effect: an elderly-word prime and a neutral-word prime, followed in both cases by the same neutral judgment task. In both studies, participants exposed to elderly words walked slightly more slowly, but the effects were small, d = .25 and d = .29, and neither difference was statistically significant. The two control conditions did not differ significantly in either experiment.

These results look very different from Bargh et al.’s original estimates of d = 1.04 and d = .79. The direction was the same, but the magnitude had fallen dramatically. Viewed as replications of the elderly-priming effect, both studies therefore provided evidence that the original effect size was substantially exaggerated.

But replication of the Bargh effect was not the main question in Dijksterhuis et al.’s article. The new theoretical hypothesis concerned behavioral contrast. After participants were primed with elderly words, some completed an additional judgment task about a specific elderly exemplar, the Dutch Queen Mother, Princess Juliana. These participants subsequently walked faster than participants in both the neutral control condition and the ordinary elderly-prime condition. This predicted contrast effect was statistically significant in both Studies 2a and 2b and became the central result of the article.

Importantly, the authors did notice that their ordinary elderly-prime condition had failed to reproduce Bargh et al.’s large assimilation effect. They described the absence of behavioral assimilation in both studies as “somewhat surprising.” However, they proposed that the neutral judgment task inserted between the prime and the walking measure may have wiped out the priming effect. Because the means were nevertheless in the predicted direction, they interpreted them as possible evidence of “residual assimilation.”

Thus, the replication failure was not literally overlooked. It was theoretically reinterpreted and became peripheral to the new finding. In the structure of the experiment, the conditions corresponding most closely to Bargh’s original elderly-prime comparison functioned merely as control conditions for demonstrating the novel contrast effect. The striking result was that thinking about a specific elderly exemplar reversed the expected behavioral effect, not that the ordinary elderly prime produced only a tiny and nonsignificant difference.

This illustrates an important consequence of focusing on statistical patterns rather than effect sizes. If the only question is whether the effect has the predicted sign, Dijksterhuis et al.’s results can be described as consistent with Bargh et al. If the question is whether the magnitude of Bargh et al.’s finding replicated, the answer is very different. An effect of approximately one standard deviation had become an effect of approximately one quarter of a standard deviation in two independent studies. By this criterion, the first evidence that the famous elderly-priming effect is not as large and robust as the original study suggested appeared fourteen years before Doyen et al. (2012).

Dijksterhuis et al.’s post-hoc explanation for the replication failure is also revealing. If inserting an ostensibly neutral judgment task can reduce an effect by roughly 75%, the effect is not very robust. The same issue applies to Bargh’s later explanation of Doyen et al.’s replication failure. If a seemingly minor difference in how participants were directed to leave the laboratory can eliminate an effect that was originally estimated at about one standard deviation, then the effect is highly sensitive to procedural details.

This distinction is important. Behavioral-priming researchers were interested in theoretically predicted moderators because they promised to specify when and why priming effects occur. But the accumulating evidence also suggested the existence of unknown moderators: seemingly incidental features of the procedure that could dramatically change the magnitude of the effect. A phenomenon that appears only under a narrow and incompletely understood set of conditions is very different from the large and robust effect suggested by Bargh et al.’s original experiments.

The second replication failure was published in 2002. Hull et al. replicated Bargh’s elderly-priming paradigm, but also measured the personality traits self-consciousness. In two studies, they found priming effects only for participants high in self-consciousness, but not for those low in self-consciousness. The average effect sizes in both studies were moderate, ds = .57, .54, but not statistically significant due to the small sample sizes in both studies.

The third replication failure was published in 2006. Cesario et al. added a condition priming youth and found that youth primes made participants walk faster, but the difference between the control condition and the elderly prime condition was small, d = .23, and not statistically significant.

In 2008, Jeffries and Fazio showed an interaction effect between elderly priming and a stopping rule on persistence on an anagram task. Importantly, priming had no main effect, d = -.05.

The first robust replication was published in 2011. N. Wyer found a strong effect of elderly priming on walking speed, d = .70, and an even stronger one on working memory, d = 1.03.

In the same year as Doyen’s article reported a failure, another article also reported a strong replication in one Study (Study 2, ds = 1.26, 1.05), but weak effects in Study 3 (d = .29, .31).

So, Doyen’s results are not remarkable in the context of other elderly priming studies. Some show strong effects, some show weak effects. The effect seems to be moderated by known and unknown factors that can contribute to variation across studies. Another factor is plain sampling error. In small samples, random variation across participants has a large effect on behavior, leading to imprecise estimates of the true population effect size.

A Meta-Analysis

There have been numerous meta-analyses of priming studies, but they typically combine a wide range of studies using different types of primes and outcome variables. A key problem with these meta-analyses is the large heterogeneity in effect-size estimates. Consequently, they provide limited information about the effect size of a specific paradigm such as elderly priming.

I conducted a meta-analysis of 15 tests of elderly priming from 12 studies reported in 8 articles. Given clear evidence of publication bias in the broader literature, I used a selection model that estimates publication bias and provides bias-corrected effect-size estimates (Vevea & Hedges, 1995). Clustered bootstrapping was used to account for the nesting of multiple tests within articles and to obtain confidence intervals and a 95% prediction interval.

The bias-corrected average effect-size estimate was d = .42, 95% CI [-.01, .71]. Although the point estimate is similar to that reported in the much larger meta-analysis by Dai et al. (2023), the confidence interval is wide and includes zero. The individual studies were generally small and therefore provided imprecise estimates of the effect. The amount of heterogeneity was also highly uncertain. The point estimate was tau = .26, but sampling error was considerabley, 95%CI [.00, .37]. The data therefore provide no compelling evidence that the discrepant findings of Bargh and Doyen reflect different true effect sizes.

This uncertainty is also reflected in the 95% prediction interval, which ranges from d = -.10 to d = 1.03 because uncertainty about the average population effect size is compounded with uncertainty about heterogeneity across studies. Thus, even a very large replication study in which sampling error is practically negligible could produce an effect ranging from essentially zero to a very large effect because increasing sample size cannot eliminate between-study heterogeneity. The available evidence therefore provides little support for a meaningful effect in the opposite direction, but it also does not tell us with much precision how strongly elderly primes may slow walking.

The meta-analysis also does not explain why Bargh obtained two significant results in studies with very small samples (N = 30). If d = .427 were the true effect, the probability of obtaining a significant result with N = 30 would be only about 20%. The probability of obtaining significant results in two independent studies would therefore be approximately 4%. Of course, it remains possible that unknown moderators produced larger effects in Bargh’s studies.

In short, a meta-analysis of the available studies in 2012 would have shown Bargh that Doyen et al.’s findings were entirely consistent with the broader evidence, even if they appeared inconsistent with his seminal results. There would have been no need to attribute the failed replication to incompetence. His more consequential mistake was to assume that his original studies and conceptual replication studies with other primes had provided clear and conclusive evidence for a general effect of elderly priming on behavior.

The Fall-Out

Concerns about the credibility of published findings in social psychology were already growing at the time. The psychologist Daniel Kahneman believed in priming effects and had featured this research prominently in his bestselling book Thinking, Fast and Slow. He was nevertheless becoming concerned about the reputation of priming research and discussed these concerns with Bargh (Bartlett, 2012). In a widely circulated email addressed to Bargh and other leading priming researchers, Kahneman proposed a straightforward way to restore confidence in the field. If Doyen et al.’s findings were a fluke or the result of experimental mistakes, Bargh and other experienced priming researchers could conduct rigorous replication studies and demonstrate that the effects were reproducible.

For example, with an effect size of d = .40, researchers could pool their resources and collect data from N = 400 participants, providing approximately 98% power to obtain a significant result. Yet Kahneman’s proposed collaborative replication effort never materialized. Instead, other researchers attempted preregistered replications of priming effects, often with disappointing results. In Dai et al.’s (2023) meta-analysis, preregistered studies produced essentially no average priming effect (d = .02), and most of these studies were replications.

In 2017, Kahneman reflected again on the problems in priming research by citing Overall (1969), who had pointed out that the prevalence of studies deficient in statistical power is not merely wasteful but potentially pernicious because it can produce a high proportion of invalid rejections of the null hypothesis among published findings. Studies such as Bargh’s elderly-priming experiments, with only 15 participants in each experimental condition, had little power to detect plausible effects and cannot explain success rates approaching 90% in published priming articles. Hundreds of published studies created the appearance of a robust phenomenon, but low power, publication bias, heterogeneity, and the near-zero effects in preregistered studies undermine the conclusion that behavioral priming was a well-established and broadly generalizable phenomenon.

Rather than joining forces to establish which of their original findings could be reproduced under well-controlled conditions, priming researchers largely moved on. Kahneman’s proposed coordinated replication effort went nowhere, and the classic behavioral-priming paradigms that had occupied such a prominent position in social psychology largely disappeared from the research frontier. In this sense, behavioral priming is practically dead. Its history provides a stark demonstration of the dangerous combination of strong beliefs and weak evidence. Researchers can find reasons to dismiss failures and construct increasingly elaborate theories to explain inconsistent findings without first establishing the robustness of the phenomenon that those theories are intended to explain. As the preceding meta-analysis shows, even the evidence available in 2012 would have revealed substantial uncertainty and would have justified Kahneman’s call for rigorous replication studies.

An obvious question, then, is why a true believer would not rise to the occasion and use the considerable resources available to a Yale professor to demonstrate that the effect could be reproduced. Why did Bargh not do it?

In an interview, Bartlett ((2013) asked Bargh essentially the same question:

So why not do an actual examination? Set up the same experiments again, with additional safeguards. It wouldn’t be terribly costly. No need for a grant to get undergraduates to unscramble sentences and stroll down a hallway.

Bargh showed little enthusiasm for returning to the laboratory simply to demonstrate that he could reproduce his original findings:

he wouldn’t want to force his graduate students, already worried about their job prospects, to spend time on research that carries a stigma. Also, he is aware that some critics believe he’s been pulling tricks, that he has a “special touch” when it comes to priming, a comment that sounds like a compliment but isn’t. “I don’t think anyone would believe me,” he says.

The same article describes an interview with Axel Cleeremans, one of Doyen’s collaborators, about their reaction to Bargh’s blog post:

It was obvious that he was so dismissive, it was close to frankly insulting. He described us as amateur experimentalists, which everyone knows we are not.” Nor did they feel that his critique of their methods was valid. Even so, they tried the experiment again, taking into account Bargh’s concerns. It still didn’t work.

Bargh’s response to Doyen et al. nevertheless expressed considerable confidence that priming was real and that their failed replication was an uninformative anomaly. He believed that the long-term future of priming research was secure because the phenomenon rested on a large and converging empirical foundation:

Like many scientists, I take the long view and continue to have faith in science as a cumulative process, with particular studies that vary in their methods and approaches converging on deeper underlying principles and mechanisms. No single experiment, standing by itself, should be the basis for concluding anything, maybe most especially in psychological research. And when a single study does not replicate another one whose findings are solidly embedded in theories of more than one scientific field and which is consistent with dozens if not hundreds of other conceptual replications, then responsible scientists—and responsible science journalists—do not rush to judgment and make claims that the entire phenomenon in question is illusory. Thomas Kuhn famously argued in The Structure of Scientific Revolutions that it is the accumulation of contrary evidence over time, not single studies, which is needed to overturn established concepts and principles in any branch of science.

A decade later, Bargh’s faith in cumulative science has been vindicated, but not in the way he expected. Replication failure after replication failure, combined with the absence of convincing new evidence from the original researchers, has shown that many claims about unconscious behavioral priming are no more securely established scientifically than Freud’s speculations about unconscious influences on behavior. The central problem was that priming researchers disproportionately published and shared their successes, creating the appearance of an unusually robust phenomenon. As shown here, even those significant findings did not tell a clean story and provided no clear evidence that elderly priming had ever produced a reliable effect.

Bargh was therefore right about one thing: no single failed replication should have settled the issue. What mattered was the accumulation of evidence over time. That accumulation simply turned out to favor Doyen’s skepticism rather than Bargh’s confidence. Maybe, deep down in his unconscious, Bargh realizes this too—which might explain why he never attempted to replicate his 1996 findings.

The Lesson

Bargh is not the first scientist to devote much of his career to an idea that ultimately failed to live up to its initial promise. This is an occupational hazard of working at the cutting edge of science. Novel research questions are difficult, methods are imperfect, and many technical problems have to be solved before a genuine breakthrough can be distinguished from a false lead. Strong belief in an idea can provide the motivation needed to overcome these obstacles, but it also creates a danger: it becomes easier to miss the warning signs provided by failed studies.

The danger increases after a theory has been published and receives attention and acclaim. At that point, researchers are no longer merely testing an interesting idea. Their reputation and professional identity may become tied to its success. Failures can then be attributed to methodological problems, incompetent experimenters, unknown moderators, or subtle boundary conditions, while successes are interpreted as evidence that the theory is correct. As the Nobel laureate Richard Feynman famously observed:

“The first principle is that you must not fool yourself, and you are the easiest person to fool.”

Young scientists can reduce this risk by learning from the mistakes of previous generations. It is often said that each generation stands on the shoulders of giants who made major discoveries. What is mentioned less often is that many equally talented scientists climbed just as high and pursued ideas that turned out to be wrong. The scientists whose theories survived were not necessarily wiser or more rigorous at the beginning; to some extent, they were fortunate that nature cooperated with their ideas. Darwin’s central theory of evolution by natural selection survived increasingly stringent tests, whereas Lamarck’s proposed mechanism of inheritance did not. Neither could have known the eventual outcome when they began trying to understand how species change.

The lesson is therefore not to avoid strong theories or strong beliefs. Science needs researchers who are willing to pursue difficult ideas for years despite setbacks. The lesson is to remain willing to discover that the idea is wrong. The stronger the personal investment in a theory becomes, the more important it is to design studies that can produce an outcome that the researcher does not want to see.

Leave a Reply