Replicability Index: A Blog by Dr. Ulrich Schimmack

Blogging about statistical power, replicability, and the credibility of statistical results in psychology journals since 2014. Home of z-curve, a method to examine the credibility of published statistical results.

Show your support for open, independent, and trustworthy examination of psychological science by getting a free subscription. Register here.

For generalization, psychologists must finally rely, as has been done in all the older sciences, on replication” (Cohen, 1994).

DEFINITION OF REPLICABILITYIn empirical studies with sampling error, replicability refers to the probability of a study with a significant result to produce a significant result again in an exact replication study of the first study using the same sample size and significance criterion (Schimmack, 2017). 

See Reference List at the end for peer-reviewed publications.

Mission Statement

The purpose of the R-Index blog is to increase the replicability of published results in psychological science and to alert consumers of psychological research about problems in published articles.

To evaluate the credibility or “incredibility” of published research, my colleagues and I developed several statistical tools such as the Incredibility Test (Schimmack, 2012); the Test of Insufficient Variance (Schimmack, 2014), and z-curve (Version 1.0; Brunner & Schimmack, 2020; Version 2.0, Bartos & Schimmack, 2021). 

I have used these tools to demonstrate that several claims in psychological articles are incredible (a.k.a., untrustworthy), starting with Bem’s (2011) outlandish claims of time-reversed causal pre-cognition (Schimmack, 2012). This article triggered a crisis of confidence in the credibility of psychology as a science. 

Over the past decade it has become clear that many other seemingly robust findings are also highly questionable. For example, I showed that many claims in Nobel Laureate Daniel Kahneman’s book “Thinking: Fast and Slow” are based on shaky foundations (Schimmack, 2020).  An entire book on unconscious priming effects, by John Bargh, also ignores replication failures and lacks credible evidence (Schimmack, 2017).  The hypothesis that willpower is fueled by blood glucose and easily depleted is also not supported by empirical evidence (Schimmack, 2016). In general, many claims in social psychology are questionable and require new evidence to be considered scientific (Schimmack, 2020).  

Each year I post new information about the replicability of research in 120 Psychology Journals (Schimmack, 2021).  I also started providing information about the replicability of individual researchers and provide guidelines how to evaluate their published findings (Schimmack, 2021). 

Replication is essential for an empirical science, but it is not sufficient. Psychology also has a validation crisis (Schimmack, 2021).  That is, measures are often used before it has been demonstrate how well they measure something. For example, psychologists have claimed that they can measure individuals’ unconscious evaluations, but there is no evidence that unconscious evaluations even exist (Schimmack, 2021a, 2021b). 

If you are interested in my story how I ended up becoming a meta-critic of psychological science, you can read it here (my journey). 

References

Brunner, J., & Schimmack, U. (2020). Estimating population mean power under conditions of heterogeneity and selection for significance. Meta-Psychology, 4, MP.2018.874, 1-22
https://doi.org/10.15626/MP.2018.874

Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17, 551–566
http://dx.doi.org/10.1037/a0029487

Schimmack, U. (2020). A meta-psychological perspective on the decade of replication failures in social psychology. Canadian Psychology/Psychologie canadienne, 61(4), 364–376. 
https://doi.org/10.1037/cap0000246

Mastodon

Primed for Equivocation

“Priming exercise” and psychological priming share a label, not necessarily a mechanism. Deliberate preparation for a known future competition provides no obvious need for unconscious goal activation, and the article presents no evidence that such activation actually occurs. The shared terminology leaves the proposed explanation primed for equivocation.

Holmberg and Kelly (2026) provide a useful critique of physiological explanations for “priming exercise,” but their attempt to connect this literature to psychological priming introduces a much less plausible mechanism. The connection appears to arise largely because the two literatures happen to use the same word. That shared terminology leaves the argument primed for equivocation.

In sports physiology, “priming exercise” refers to a deliberately performed bout of exercise intended to improve performance later that day. In cognitive and social psychology, priming refers to prior exposure to a stimulus that subsequently alters processing or behavior, sometimes without awareness or conscious intention. These are fundamentally different uses of the term. The fact that both involve something occurring before something else does not imply that they share a psychological mechanism.

This distinction becomes especially important when the authors invoke nonconscious goal priming. They suggest that exercise might activate performance goals and related behavioral representations that persist until later testing. They cite classic social-psychological priming research, including Bargh and colleagues, to support the possibility that goals can be activated without awareness and subsequently influence behavior.

But the proposed mechanism is poorly matched to the phenomenon being explained. An athlete does not ordinarily encounter a “priming exercise” incidentally. The athlete performs it because a competition or performance test is coming later. The later performance goal is therefore likely to have been activated before the exercise begins:

competition later → intention to prepare → priming exercise.

The goal is not plausibly dormant until the exercise somehow activates it unconsciously. Indeed, the goal is probably one of the reasons the athlete performs the exercise in the first place. During the exercise the athlete may also consciously think about the competition, technique, pacing, readiness, or expected benefits. Under those circumstances, invoking nonconscious goal activation is not merely unnecessary; the proposed causal sequence is almost backwards.

The authors’ own examples illustrate the problem. They discuss athletes rehearsing particular pacing strategies, using metronomes, receiving verbal cues, believing that squatting will improve subsequent jumping, and developing confidence or expectations about later performance. These are readily understood as deliberate preparation, task practice, expectancy, motivation, or attentional effects. None requires a nonconscious priming mechanism.

The scientific evidence offered for the nonconscious account is also weak. The article provides no exercise experiment demonstrating that a priming-exercise bout activates a previously inactive goal outside awareness and that this activation subsequently causes improved performance hours later. In fact, the authors repeatedly acknowledge that these possibilities “may” occur, are “hypothesized,” or “have yet to be directly examined.” Thus, the proposed mechanism is not an empirical finding from the exercise literature.

Instead, support is imported from a different literature on behavioral and goal priming. That literature is itself scientifically controversial, and citing classic demonstrations does not establish that the same mechanism operates in an entirely different situation involving intentional athletic preparation. The conceptual inference appears to be:

psychology calls something “priming”

  • exercise science calls something “priming”
    → psychological priming may explain exercise priming.

But identical terminology is not evidence of mechanistic equivalence.

The distinction matters because the paper already identifies much more plausible explanations. Task-specific practice, motor learning, expectancy, researcher effects, motivation, and explicit performance preparation could all produce later performance changes. These mechanisms fit the actual structure of the situation: athletes know that performance is coming and intentionally prepare for it. The nonconscious goal-priming hypothesis adds an unnecessary and poorly supported causal layer.

The problem can therefore be summarized simply: an implausible mechanism is invoked to explain effects that have not been shown to require that mechanism, and the empirical justification comes largely from a separate literature that happens to use the same word.

Methodological Problems in Claims About “Conscious” and “Preconscious” Influencer Effects

Mir, I. A. (2026). Influencer’s physical attractiveness and content aesthetics: Conscious and preconscious determinants of fashion-branded content engagement on Instagram. Journal of Creative Communications, 21(2), 201–219. https://doi.org/10.1177/09732586241288672

“The authors invoke an unproven perception–behavior mechanism to explain causation that was never observed. A direct path in a cross-sectional SEM establishes neither causation nor preconscious processing.”

Mir (2026) examines whether fashion influencers’ physical attractiveness and the aesthetics of their branded content predict followers’ engagement on Instagram. The study uses survey responses from 300 followers of 15 fashion influencers in Pakistan and analyzes the proposed relationships with structural equation modeling and mediation analyses. The main empirical finding is straightforward: followers who rate influencers and their content more positively also report more favorable attitudes and greater engagement. The methodological problem is that the article draws causal and psychological-process conclusions that the design cannot support.

The most serious problem is that all variables were measured in a single cross-sectional self-report survey. Participants simultaneously reported how attractive they considered the influencer, how aesthetically pleasing they considered the content, their attitude toward that content, and how often they viewed, liked, commented on, and shared it. Nothing was manipulated, and there was no temporal ordering of the variables. Nevertheless, the article repeatedly describes attractiveness and aesthetics as factors that “cause,” “trigger,” “stimulate,” or “activate” engagement. Those causal statements do not follow from the design.

For example, the proposed model assumes

attractiveness → attitude → engagement.

But the same covariance pattern is compatible with numerous alternatives. Followers who engage frequently with an influencer may develop more positive attitudes and subsequently rate that influencer as more attractive. A general liking or identification with the influencer could simultaneously increase attractiveness ratings, content-aesthetic ratings, attitudes, and engagement. The structural equation model cannot distinguish among these explanations.

The sampling procedure makes this problem particularly important. Participants had to have followed one of the selected influencers for more than six months. Thus, the sample is already conditioned on sustained interest in the influencer. People who disliked the influencer, found the content unattractive, or disengaged from it are systematically less likely to appear in the sample. The resulting correlations describe differences among an already selected group of followers; they provide weak evidence for the article’s practical recommendation that firms should hire physically attractive influencers because attractiveness causes engagement.

The article’s central methodological error is even more fundamental. It claims to distinguish a “conscious” route from a “preconscious” route. The mediated path

attractiveness/aesthetics → attitude → engagement

is interpreted as conscious influence, whereas a remaining direct path from attractiveness or aesthetics to engagement is interpreted as evidence of a preconscious perception–behavior process.

A direct regression coefficient is not a measure of unconscious processing.

If attractiveness predicts engagement after statistical adjustment for an explicit attitude measure, this merely shows residual covariance between those variables. That residual association could reflect measurement error in attitude, omitted mediators, stable preferences, common response tendencies, reverse causation, or numerous other processes. Nothing in the study measures awareness, intention, automaticity, processing speed, or participants’ ability to report the causes of their behavior. Consequently, the data provide no evidence that engagement was “preconscious” or unintentional.

For the same reason, the indirect path through an explicit attitude measure does not establish a conscious causal mechanism. Participants consciously completed the attitude questionnaire, but that does not mean the psychological process producing their engagement operated consciously. Statistical mediation and conscious psychological mediation are different concepts.

This problem is especially consequential because the “preconscious” interpretation is one of the article’s principal theoretical contributions. The authors explicitly invoke the perception–behavior literature, including Bargh et al. (1996), to justify the claim that a significant direct path demonstrates automatic behavior. But the current research contains none of the experimental procedures that would be required to test an automatic perception–behavior effect. The SEM therefore cannot adjudicate between conscious and unconscious processes.

The measures themselves also create substantial interpretive problems. On page 209, physical attractiveness is measured with “stylish,” “good looking,” “sexy,” and “elegant.” Content aesthetics is measured with “striking,” “wonderful,” “fascinating,” and “lovely,” while attitude is measured with “pleasant,” “good,” “likeable,” and “my favourite.” These constructs are conceptually and evaluatively intertwined. “Wonderful” and “lovely,” for example, are not narrowly aesthetic judgments, while “stylish” and “elegant” are not purely measures of physical attractiveness. Much of the model may therefore reflect a broad positive-evaluation factor rather than distinct psychological constructs connected by causal pathways.

The observed correlations are consistent with this concern. Physical attractiveness correlates .64 with content aesthetics and .64 with attitude; content aesthetics correlates .66 with attitude and .68 with reported content consumption. Demonstrating discriminant validity with the Fornell–Larcker criterion does not eliminate the possibility that halo effects or general liking strongly influence all of these ratings.

The attempt to dismiss common-method bias is also inadequate. All predictors, mediator variables, and outcomes were obtained from the same respondent at the same time using similar rating formats. The authors test whether several sets of items can be represented by a single latent factor and conclude that poor single-factor fit shows that common-method bias is not important. That conclusion does not follow. Common-method variance does not require every item to load on one factor. Several distinguishable constructs can coexist while correlations among them are inflated by shared method, evaluative consistency, acquiescence, or halo effects.

The proposed “snowball effect” suffers from the same causal problem. The authors find that self-reported consumption behaviors—viewing, reading comments, and liking—predict contribution behaviors such as commenting and sharing, and conclude that consumption gradually causes users to progress toward more active participation. Yet consumption and contribution were measured simultaneously. Someone who frequently comments and shares influencer content almost necessarily also views and consumes that content. A positive cross-sectional association therefore does not demonstrate a temporal progression from passive to active engagement. Testing a snowball process would require longitudinal evidence showing that earlier consumption predicts subsequent increases in contribution.

There may also be an unmodeled dependence problem. The 300 respondents followed one of only 15 macro-influencers. Followers of the same influencer are not necessarily independent observations. Influencers may differ systematically in appearance, production quality, follower demographics, posting frequency, and baseline engagement. Those influencer-level characteristics could generate correlations among respondent ratings. The reported SEM appears to treat all 300 followers as independent rather than accounting for clustering by influencer.

Another reporting issue concerns the engagement scale. The response categories are described as 1 = “very often,” 2 = “often,” 3 = “sometimes,” and 4 = “never.” Thus, larger numerical values indicate less engagement. Yet positive path coefficients are consistently interpreted as greater attractiveness and aesthetics producing greater engagement. The article does not clearly state in the reported method that these scores were reverse-coded. If they were reversed before analysis, that transformation should have been explicitly documented. If they were not, the substantive interpretation of the coefficients would be reversed.

Finally, the statistical success of the model should not be confused with strong evidence for the hypotheses. All five proposed hypotheses are supported, including the weakest direct attractiveness effect, β = .12, t = 2.19. The study was not preregistered, and the sample-size justification consists largely of the statement that N = 300 is sufficient for purposive sampling and structural equation modeling rather than an a priori power analysis tied to the focal effects. Good model-fit indices demonstrate that a specified covariance model can reproduce the observed covariance matrix; they do not establish that the arrows in Figure 2 represent the true causal processes.

The study therefore supports a much narrower conclusion than the article claims. Among existing long-term followers of fashion influencers, positive ratings of influencer attractiveness and content aesthetics are associated with positive attitudes and greater self-reported engagement. That descriptive association is plausible and potentially useful.

The study does not establish that physical attractiveness or content aesthetics cause engagement, that attitude mediates these effects causally, that any residual direct relationship reflects a preconscious perception–behavior mechanism, or that passive engagement develops over time into active contribution.

A suitable experimental design would manipulate influencer attractiveness and content aesthetics independently, randomly assign participants to conditions, measure actual engagement behavior, and include measures capable of testing awareness or automaticity. A longitudinal design would be required to test the proposed consumption-to-contribution “snowball” process. Without such evidence, the article’s strongest psychological claims are interpretations imposed on cross-sectional correlations rather than findings produced by the research design.

Auditing a Poor Audit of Elderly Priming

Costa, T. (2026). The Bayesian audit: Evaluating the proportionality of scientific claims to evidence—a case study on social priming and walking speed. Frontiers in Psychology, 17, 1799078. DOI: 10.3389/fpsyg.2026.1799078


This article applies Bayes’ theorem to one conveniently selected t value and calls the result a ‘Bayesian audit.’ Most readers may simply ignore it because it appeared in Frontiers in Psychology. Those who want a more substantive reason can point to this review.

Costa (2026) introduces a “Bayesian audit,” a six-step framework intended to evaluate whether the strength of scientific claims is proportional to the evidence supporting them. The idea is sensible. Statistical significance does not tell us how strongly we should believe a scientific claim, and surprising claims based on weak evidence deserve particularly careful scrutiny. Costa illustrates the proposed method with one of social psychology’s most famous findings: Bargh, Chen, and Burrows’s (1996) claim that priming college students with words related to old age caused them to walk more slowly afterward.

Unfortunately, the audit itself is problematic. It misrepresents important features of the original study, considers only a fraction of the available evidence, and reduces a question about the magnitude and robustness of an effect to a comparison between a null and an inadequately specified alternative hypothesis.

The first problem is surprisingly basic. Costa describes the original finding as based on a study with approximately t(28) = 2.0 and p ≈ .05. But Bargh et al. actually reported two elderly-priming experiments. In Experiment 2a, the comparison was t(28) = 2.86, p < .01. They then conducted Experiment 2b as a replication and again reported slower walking, t(28) = 2.16, p < .05. Costa appears to approximate the weaker second result while failing to mention the stronger first result or even that the original article contained two studies.

That is an odd starting point for an audit. If the purpose is to reconstruct how much evidence supported the claim in 1996, both original studies should be included.

Costa also incorrectly describes participants as being “subliminally exposed to words related to old age.” They were not. Participants consciously read words while completing a scrambled-sentence task. The claimed unconscious component was that participants supposedly did not realize that the elderly-related words subsequently affected their walking. Bargh et al. themselves explicitly distinguished this procedure from subliminal priming; Experiment 3 of their paper used genuinely subliminal presentation of faces.

This distinction matters because Costa uses the apparent implausibility of unconscious effects on motor behavior to motivate skeptical prior probabilities. One should at least characterize the causal claim correctly before assigning a prior to it.

The treatment of replication evidence is even more problematic.

Costa cites Doyen et al. (2012) and Harris et al. (2013) as subsequent replication attempts. Doyen et al. did replicate the elderly-walking paradigm. In a substantially larger study using automated measurement, they found essentially no priming effect. Their second experiment further suggested that experimenter expectations could influence the result.

Harris et al. (2013), however, did not replicate elderly priming at all. They attempted to replicate Bargh et al.’s 2001 high-performance goal-priming experiments, in which achievement words were supposed to improve performance on a cognitive task. Calling Harris et al. a replication of the elderly-walking finding is simply an error.

More importantly, why is a Bayesian audit conducted in 2026 based primarily on one t statistic from 1996?

There is now a substantial literature on behavioral priming. Dai et al. (2023), for example, meta-analyzed 351 studies and 862 effect sizes and concluded that behavioral priming effects could be detected across a large literature. Conversely, Mac Giolla et al. (2024) examined 70 close replication attempts of 49 social-priming findings. Ninety-four percent produced smaller effects than the originals, only 17% were significant in the predicted direction, and among 52 replications conducted without an original author, none was significant in the original direction; the pooled effect for those independent replications was essentially zero.

These sources do not necessarily settle the question. Meta-analyses themselves can be distorted by publication bias and other forms of selection. But that is precisely why an audit should examine them critically. An audit of a 30-year-old scientific claim should evaluate the accumulated evidence, not simply convert one selected original result into a Bayes factor.

There is an even more fundamental problem with the statistical question Costa asks.

Costa assigns prior probabilities of .05, .10, and .20 to the alternative hypothesis and combines these with an estimated Bayes factor of approximately 3. This yields posterior probabilities of .14, .25, and .43, respectively. The arithmetic is straightforward. The interpretation is not.

Why should the prior probability that the effect exists be .05 or .10?

Costa acknowledges that these values are illustrative rather than derived from an elicitation procedure. But these priors largely determine the conclusion that posterior belief remains low. Starting with a 5% probability and multiplying the prior odds by a Bayes factor of 3 inevitably produces a low posterior probability.

More importantly, what exactly is the hypothesis whose prior probability is 5%?

There is a major difference between these propositions:

elderly-related words have exactly zero effect on walking speed;

elderly-related words have some nonzero effect;

elderly-related words have a psychologically meaningful effect;

elderly priming produces effects of the magnitude originally reported;

automatic stereotype activation reliably produces consequential behavioral changes.

These are not the same hypothesis.

The scientifically interesting issue today is probably not whether the population effect is exactly zero. The effect could be d = .05 or d = .10. Such an effect would make the point null hypothesis technically false while providing little support for the dramatic theoretical interpretation of the original experiments.

This is why effect sizes matter. Bargh et al.’s original studies implied very large effects. Subsequent evidence raises the possibility that the true effect, if it exists at all, is much smaller. A useful audit therefore needs to ask how large the effect is and how precisely it has been estimated—not merely whether H0 or H1 receives the larger Bayes factor.

Costa’s procedure also conflates two different kinds of priors. One is the prior model probability: how likely H1 is relative to H0 before seeing the data. The other is the prior distribution over possible effect sizes within H1. A Bayes factor for a composite alternative necessarily depends on the latter. Yet the article emphasizes sensitivity to prior model probabilities while giving much less attention to the effect-size assumptions used to obtain BF₁₀ ≈ 3.

This creates another problem with Costa’s distinction between “evidence” and “belief.” He describes the Bayes factor as quantifying evidence supplied by the data and posterior probability as combining this evidence with prior belief. But a Bayes factor is not simply a property of the observed data. It depends on the statistical models being compared, including the distribution of effect sizes assumed under the alternative hypothesis.

There is also an internal inconsistency in the treatment of the replication evidence. Costa states that the later replication attempts yielded Bayes factors close to 1 and therefore had little evidential impact. That makes sense: a Bayes factor of 1 leaves prior odds unchanged. Yet the subsequent synthesis says that posterior belief “collapses under replication.” It cannot do both. Replications with BF ≈ 1 cannot cause posterior belief to collapse. To demonstrate such a decline, one would need Bayes factors favoring the null or another competing model and then accumulate this evidence formally.

Publication bias is another conspicuous omission. The article is motivated by the replication crisis and explicitly acknowledges that biased data limit the usefulness of evidential measures. Yet the actual Bayesian calculation treats the published Bargh result as though it were an observation selected independently of statistical significance.

That is unrealistic. A BF of 3 obtained from a randomly selected study and a BF of 3 obtained from a literature in which statistically significant and theoretically exciting findings were preferentially published do not have the same evidential implications. If selection contributed to the replication crisis, an audit of the original published evidence needs to take selection seriously.

Costa also describes the original study’s low statistical power as an additional reason for skepticism. Low power certainly matters because significant results from low-powered studies tend to exaggerate effect sizes, particularly in a selected literature. But sample size has already entered the likelihood used to compute the Bayes factor. Low power is therefore not independent evidence against the hypothesis. The additional concern arises from selection, analytic flexibility, measurement error, and effect-size inflation.

The deeper problem is that Costa reduces a scientific question to H0 versus H1 when several competing explanations exist. Doyen et al.’s work raised experimenter expectancy as one possible explanation. Other possibilities include a genuinely small priming effect, effects restricted to particular conditions or individuals, procedural artifacts, or some combination of these mechanisms. A Bayes factor contrasting an exact-zero model with a generic nonzero-effect model cannot determine which causal explanation is correct.

This is especially important because rejecting H0 would not establish Bargh’s theory. Even convincing evidence for a tiny difference in walking speed would not demonstrate that automatic stereotype activation generally controls overt behavior.

The Bayesian audit is therefore based on a reasonable principle but a poor demonstration. Scientific claims should indeed be proportional to evidence. The problem is that assessing proportionality requires accurately identifying the original evidence, considering the accumulated replication literature, evaluating publication bias, distinguishing statistical from substantive hypotheses, and estimating plausible effect sizes and their uncertainty.

Ironically, the elderly-priming case illustrates the weakness of Costa’s audit more effectively than it illustrates its strengths. A proper audit should not ask merely whether one selected t statistic changes the odds that an effect is exactly zero. It should ask what three decades of evidence tell us about the magnitude, robustness, boundary conditions, and causal interpretation of the phenomenon.

On those questions, uncertainty remains. There may be a small elderly-priming effect. The evidence does not establish that the effect is exactly zero. But neither does the accumulated evidence support taking the spectacular effects reported in 1996 at face value. The important scientific task is to estimate what effect remains after accounting for uncertainty and bias. That requires more than Bayes’ theorem applied to one conveniently chosen t value.

Old Evidence for a Fragile Priming Theory

Przybylinski, E. (2026). Whatever you say: Changing transference-based problem behavior with if–then plans. Self and Identity. Advance online publication. https://doi.org/10.1080/15298868.2026.2613846


Przybylinski (2026) reports two experiments examining whether implementation intentions can prevent problematic behaviors triggered by transference. The theoretical logic is straightforward: subtle resemblance to a significant other is assumed to activate that person’s representation automatically, which can then influence memory, goals, and behavior. An if–then plan is proposed to prevent the activated representation from guiding subsequent behavior.

The principal concern is the credibility of the evidence on which this argument rests.

The article treats automatic behavioral priming as a well-established foundation. For example, it cites Bargh et al. (1996), Bargh et al. (2001), Chartrand and Bargh (1999), and related studies as evidence that contextual cues can automatically trigger overt behavior without awareness or intention. Yet behavioral priming is precisely one of the areas most affected by the replication crisis. The article does not discuss this change in evidential status. Thus, evidence that was considered persuasive when these studies were conducted is largely presented in 2026 as though subsequent replication failures had not occurred.

This matters because the two experiments are themselves products of that earlier research era. The author explicitly states that they were conducted as dissertation research in the late 2000s, before preregistration became standard. Study 1 included only 60 participants, or 20 participants in each of the three strategy conditions, despite testing interaction hypotheses. Study 2 included 47 participants. The manuscript provides a power justification, but because the studies were not preregistered, it is unclear whether the reported power analysis reflects a prospectively specified design decision. The reported “post hoc power” provides little additional information.

The results are remarkably successful. The focal interactions involving behavioral readiness, memory, and overt submissive behavior are all statistically significant in the predicted direction. The reported effects are also very large, frequently exceeding d = 1 and reaching d = 1.69 in Study 1. Across the major focal tests, the success rate is effectively 100%, while average observed power based on the reported effects is roughly 80%.

A perfect success rate is not impossible when power is 80%, but it is more successful than expected. Schimmack (2012) emphasized that unusually high success rates relative to estimated power can indicate that published effect sizes and success rates should not be taken at face value. Here, the outcomes within each experiment are dependent, so a simple excess-success or incredibility calculation would not be appropriate. With only two studies, there is no statistical smoking gun. Nevertheless, the combination of small samples, large effects, multiple significant focal outcomes, and absence of preregistration warrants substantial caution.

The historical timing makes this concern more important. These experiments were conducted approximately 15 years before their publication. During that interval, psychology experienced a replication crisis that directly challenged the credibility of the behavioral-priming literature on which the article relies. Yet the 2026 article does not supplement the old experiments with a contemporary, adequately powered, preregistered replication.

This is particularly striking because the central experiment is readily replicable. The author remains at the same institution where the original research was conducted, and the paradigm requires undergraduate participants rather than an unusually difficult population. A preregistered replication with a substantially larger sample could have provided highly informative evidence about whether the large effects observed in the original dissertation studies survive contemporary scrutiny.

The absence of such a replication changes how the evidence should be interpreted. The results are not invalid merely because they were collected before the replication crisis. But neither the reported effect sizes nor the perfect pattern of statistical success should be treated as reliable estimates of the underlying effects without independent replication.

A Speculative Theory of Replication Failures in Social Psychology: The Empty-Self Metaphor

Klein, J. W., & Swann, W. B., Jr. (2026). Social psychology’s empty-self metaphor and the replication crisis. Perspectives on Psychological Science, 21(2), 138–153. https://doi.org/10.1177/17456916251401849.

Klein and Swann (2026) offer an interesting diagnosis of the replication crisis, but the evidence does not support the strength of their theoretical interpretation.

Their central empirical observation is striking. They coded 41 hypotheses from Many Labs 1 and 2 as either consistent with an “empty-self” metaphor or not. None of the nine “empty-self” hypotheses replicated, whereas 26 of 32 “not-empty-self” hypotheses replicated, a difference of 81 percentage points. Given the small number of empty-self studies, however, the estimate is much less precise than the point estimate suggests; an approximate 95% confidence interval for the difference is about 47 to 91 percentage points. The authors appropriately describe the result as preliminary.

The more serious problem is interpretation. The analysis is correlational. Studies were not randomly assigned to use an “empty-self” theory, and the authors did not code plausible confounding variables that could themselves predict replication success. These include the sample size and statistical power of the original study, the strength of the situational manipulation, whether the manipulation was consciously perceived, the causal proximity between manipulation and outcome, and the prior plausibility of the predicted effect. Their own coding examples illustrate the problem. A gray background affecting support for austerity is coded as “empty self,” whereas paying for a workshop affecting attendance is coded as “not empty self.” The latter is still a situational effect. What differs most obviously is that payment is a strong and behaviorally relevant manipulation, whereas background color is a weak and remote one.

Thus, the empirical result may show that studies proposing large effects of weak, incidental situational manipulations replicate poorly. That is interesting, but it is not the same as showing that studies fail because they neglect an enduring self.

The theoretical explanation is also speculative and at times internally strained. Klein and Swann sometimes treat replication failures as evidence that subtle situational manipulations have little or no effect. In the Many Labs studies, this inference can be justified when very large samples produce narrow confidence intervals around zero. But this cannot be generalized to all failed replications. In other literatures, including elderly priming, the available confidence intervals may still be compatible with small effects. Failure to obtain significance is not equivalent to demonstrating an effect of exactly zero.

At the same time, Klein and Swann suggest that subtle situational effects may depend on the person. They argue that people may respond only to cues to which they are “tuned” and that primes may work only when they connect to existing self-representations. That possibility is not new to priming theory. Priming effects have long been assumed to depend on whether participants possess the relevant stereotype or representation, and some priming studies have explicitly predicted interactions—for example, effects of religious primes that depend on participants’ religiosity.

But this moderation account has an important implication. If a prime affects people for whom it is relevant and has little effect on others, the population-average effect should normally be reduced, not eliminated. Unless one assumes theoretically unusual crossover interactions in which the prime produces effects in opposite directions for different people, sufficiently large studies should still detect a small average effect. For elderly priming, for example, it is easy to imagine that some participants might be more responsive to an elderly stereotype than others. It is much harder to explain why the same prime should make another substantial group walk faster.

This creates a useful empirical question that Klein and Swann do not examine. Their 41 hypotheses should be coded not only as “empty self” or “not empty self,” but also according to whether the original hypothesis predicted a main effect or a Person × Situation interaction. If nearly all of the “empty-self” studies in Many Labs tested simple main effects, that matters because the larger priming literature already contains many moderator and interaction hypotheses. The Many Labs sample would then not represent the full theoretical range of priming research. More broadly, the authors should distinguish studies proposing universal effects from studies explicitly predicting conditional effects.

A stronger analysis would therefore code each study independently for situation strength, awareness of the manipulation, personal relevance, causal distance between manipulation and outcome, original sample size and power, prior plausibility, and whether the prediction was a main effect or an interaction. Only then could one determine whether an “empty-self” construct predicts replication after plausible alternative explanations have been taken into account.

The irony is that Klein and Swann criticize social psychology for building theories on weak evidence, but their own “empty-self” explanation is itself a speculative theory supported by weakly diagnostic data. The empirical pattern—0% versus 81% replication—is interesting. The claim that this pattern is caused by neglect of an enduring self is not established.

On a 1-to-5 scale from speculative theory to theory explaining highly credible phenomena, I would rate the article about 2/5. The phenomenon to be explained is credible: some classes of social-psychological findings replicate poorly. The proposed explanation—that they fail because social psychology adopted an “empty-self” metaphor—remains largely speculative.

A Mostly Speculative Social Identity Theory of Digital Identity

“Theories are a dime a dozen” (Ed Diener, personal communication)

Bingley, W. J., Worthy, P., Wiles, J., & Haslam, S. A. (2026). A social identity theory of digital identity. Perspectives on Psychological Science, 21(4), 346–382. https://doi.org/10.1177/17456916261419813.



Bingley and colleagues (2026) propose a new “social digital identity theory” (SDIT) to explain how social identities operate across online, offline, and hybrid environments. The theory is unusually explicit: the authors formulate 20 propositions concerning identity salience, digital platforms, embodiment, well-being, group functioning, polarization, culture, and other outcomes.

On a 1-to-5 continuum from speculative theory to theories that explain highly credible empirical phenomena (Speculative versus Explanatory Theory: A Review of “Narrative Embodiment” – Replicability-Index), I would give SDIT a rating of 2/5 (see .

This is not a judgment that the theory is uninteresting or implausible. A rating of 2 means that much of the distinctive theory remains speculative and that the empirical findings used to motivate it have not been systematically evaluated for credibility.

There are good features. Bingley et al. clearly distinguish established ideas from novel hypotheses. They explicitly acknowledge that many of their propositions—particularly P4 through P11, P18, and P20—have not previously been tested and require empirical evaluation. Other propositions are extensions of the much older social-identity and self-categorization literature. Thus, they do not claim that all 20 propositions are established facts.

The problem is different. When published studies are cited as empirical support, the authors rarely ask how strong that evidence actually is. There is no systematic consideration of statistical power, independent replication, publication bias, or whether an apparently supportive literature consists mainly of selected significant findings.

Embodiment provides a good example

Proposition 5 states that social-identity salience is shaped by embodiment. To motivate this proposition, the authors cite studies suggesting that heat activates anger concepts, social rejection makes rooms feel colder, keeping secrets feels physically burdensome, fist clenching activates related concepts, and facial-muscle activation influences experience. They then extend this reasoning to virtual embodiment and the Proteus effect.

A particularly revealing example concerns elderly priming.

The authors cite a virtual-reality study by Reinhard et al. in which participants embodied an older or younger avatar. Participants who had embodied the older avatar subsequently walked more slowly during the first part of a walking test than participants who embodied the younger avatar, p = .033. However, these differences disappeared during the second half of the walk.

Bingley et al. describe this result as operating similarly to “classic priming studies in social psychology” and cite Bargh, Chen, and Burrows (1996).

That citation is problematic because Bargh et al.’s famous claim that elderly-related primes cause people to walk more slowly became one of the best-known examples of the replication problems in social psychology. Doyen et al. (2012) conducted a larger replication using automated measurement and failed to reproduce the original effect in their first experiment.

None of this replication history is mentioned.

The result is an evidential chain that looks stronger on paper than it really is:

reported embodiment effects → elderly-avatar study → classic elderly priming → support for an embodiment proposition.

There are several citations, but citations are not the same thing as strong evidence.

The elderly-avatar study may be interesting, and the possibility that virtual embodiment affects subsequent behavior certainly deserves further study. But a marginally significant result that is then linked to a famous but poorly replicated priming finding does not provide strong evidence for a broad theoretical proposition about embodiment and social identity.

The same caution applies to meta-analytic evidence. A meta-analysis can summarize a literature very precisely while still giving a misleading estimate if the underlying literature is affected by publication bias and selective reporting. The important question is therefore not simply whether a theory can cite studies—or even a meta-analysis—that support a proposition. The question is whether the phenomenon itself has been demonstrated with credible evidence.

Why 2 rather than 1?

SDIT deserves more than the lowest score because it is not free-form speculation. It builds on established theoretical traditions, makes explicit predictions, distinguishes some novel claims from older ones, and repeatedly identifies propositions that require future empirical tests.

But it does not deserve a middle or high score because much of the distinctive theory is not yet explaining firmly established empirical phenomena. Moreover, the embodiment section demonstrates that the authors sometimes treat published findings as evidence without critically assessing whether those findings survived the replication crisis.

Thus, the 2/5 rating means:

SDIT is a structured and testable theoretical proposal with some grounding in established research, but much of its distinctive content remains speculative, and the empirical literature used to motivate some propositions is treated too uncritically.

The citation of Bargh et al. (1996) is a particularly clear example. In 2026, a theory article should not invoke elderly priming as supportive evidence without informing readers that the original phenomenon itself remains empirically uncertain.

Individual Differences Do Not Salvage The Priming Train Wreck

Aytürk, E., & Saribay, S. A. (2026). Neglect of the individual as a neglected problem: The relevance of combined idiographic-nomothetic approaches for social psychology today. Self and Identity. https://doi.org/10.1080/15298868.2026.2697941

Aytürk and Saribay (2026) make a useful methodological argument in “Neglect of the individual as a neglected problem.” Social psychology typically averages across people even though the same situation may have different psychological meanings for different individuals. They advocate greater use of personalized stimuli and intensive repeated-measures designs to identify person-specific processes before generalizing across people.

They also suggest that this problem might help explain replication failures in social priming. For example, “elderly” might evoke a frail grandparent for one person and a vigorous retiree for another. Averaging across these individuals could obscure different or even opposing effects. This is a reasonable hypothesis. It is also useful that the authors explicitly acknowledge low statistical power, questionable research practices, and publication bias as established contributors to replication failures.

The problem is that they nevertheless cite Bargh et al. (1996) rather uncritically. The famous finding that activating the elderly stereotype makes people walk more slowly is treated as an example of a potentially heterogeneous priming effect, without informing readers that the original effect became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original effect in their first experiment.

There were earlier attempts to explain the inconsistent findings by invoking individual differences. For example, Cesario et al. (2006) reported that elderly priming slowed participants with relatively positive attitudes toward elderly people but sped up those with negative attitudes. Other studies similarly reported moderator effects. But these small-sample studies do not provide strong evidence that a robust main effect was merely hidden by heterogeneity. Cesario et al.’s study, for example, had only about 67 participants for the relevant analysis and did not obtain a conventionally significant overall priming effect. The evidential burden therefore shifted to an interaction estimated from an even less powerful design.

This is important because interaction effects are generally more difficult to estimate precisely than main effects. Finding a significant personality moderator in a small study does not establish that a failed main effect was really caused by individual differences. The moderator itself needs adequate power and independent replication. A more detailed discussion of these purported replications of elderly priming is available in “Elderly Priming: Did It Ever Work?”

Here Aytürk and Saribay’s methodological proposal may actually provide a better way forward. Intensive repeated measurement can obtain much more information from each participant and therefore can study within-person Person × Situation effects with fewer participants than a conventional design may require. This comes at a cost: many observations per person are needed, and ecological or experience-sampling studies typically sacrifice some of the experimental control available in tightly controlled laboratory experiments. Moreover, any between-person moderator still ultimately depends on having enough individuals.

Thus, the proposal itself is worth pursuing. Low power means that the failed priming literature often provides lack of evidence rather than definitive evidence of absence. Small, conditional, person-specific priming effects remain possible.

But they remain hypotheses to be demonstrated. Individual heterogeneity can explain variation in a real effect; it cannot by itself establish that an unreliable effect is real. A 2026 article explicitly concerned with the replication crisis should therefore not cite Bargh et al. (1996) without acknowledging the replication failures that make the existence of the phenomenon itself uncertain.

Is Subliminal Manipulation Really the Threat?

Smolinski, J., Smolinski, R., Kesting, P., & Kröcher, F. (2026). Ethical guidelines for designing, developing, and deploying AI negotiation agents. Group Decision and Negotiation, 35, 56. https://doi.org/10.1007/s10726-026-10011-2

One of the article’s central psychological concerns is that AI negotiation agents could exploit human vulnerabilities through sublimiminal or unconscious influence. This possibility sounds alarming because it suggests that an AI agent might manipulate a negotiator without the person even realizing that an influence attempt is taking place.

But the psychological evidence cited for this possibility is surprisingly weak.

The authors cite Bargh et al. (1996), one of the best-known studies of unconscious behavioral priming. Its famous elderly-priming experiments reported that participants exposed to words associated with old age subsequently walked more slowly. Yet the article does not mention that this finding later became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original priming effect in their first experiment. The subsequent history of the finding raises further questions about whether elderly priming was ever a reliable phenomenon; I review that evidence in more detail in “Elderly Priming: Did It Ever Work?

The problem is broader than the Bargh study. Subliminal persuasion has never been shown to be a particularly powerful method for changing consequential consumer choices. An older meta-analysis of subliminal advertising estimated an effect of only r = .059 on consumer choice. Later positive studies suggest that effects depend on restrictive boundary conditions, but even these results are not robust. Thus, it remains scientifically unknown if or when subliminal attempts to manipulate people actually work.

This raises an important question for the ethical analysis. Why would an AI negotiation agent need subliminal manipulation at all?

Ordinary persuasion operates in full awareness and can be far more direct. A seller can explicitly frame an offer as prestigious, scarce, safe, innovative, or socially desirable. A brand can persuade consumers that “Apple is cool,” and a consumer may knowingly pay considerably more for an Apple product because that identity has value to them. Nothing has to be flashed below the threshold of awareness. The person can see the message, understand the message, and still be influenced by it.

That form of influence is potentially much more relevant to AI negotiation agents. An AI can learn what arguments appeal to a particular person, emphasize some information rather than other information, frame alternatives strategically, appeal to identity or social norms, create urgency, flatter, build trust, or exploit known preferences. None of these processes requires a controversial theory of unconscious behavioral priming.

This distinction matters because the paper arguably directs attention toward the more sensational but less empirically credible risk. The image of an AI system exploiting subliminal psychological mechanisms is disturbing, but current psychology provides little evidence that such techniques constitute a powerful means of controlling behavior. More ordinary forms of conscious persuasion are both more plausible and potentially more consequential.

The ethical concern is therefore legitimate, but the psychological rationale should be reformulated. Rather than claiming that AI will be able to exploit “unconscious priming far more precisely” than humans, a more defensible concern is that AI may become exceptionally effective at personalized persuasion using psychological processes that operate with the target’s awareness.

The problem with citing Bargh is that it directs attention to irrational fears about subliminal manipulation, when AI may use much more powerful strategies to manipulate people with information that is in plain sight.

Speculative versus Explanatory Theory: A Review of “Narrative Embodiment”

Shimada, S., Tanaka, S., Morioka, S., Rode, G., Roy, J.-M., & Rossetti, Y. (2026). Narrative embodiment: A conceptual framework linking the narrative self and the embodied self. New Ideas in Psychology, 83, 101285. https://doi.org/10.1016/j.newideapsych.2026.101285

Shimada and colleagues (2026) propose “narrative embodiment” as a conceptual framework linking two aspects of the self: the narrative self and the embodied self. Their central proposal is that “character,” borrowed from Paul Ricoeur, functions as an intermediary between narratives about who we are and relatively stable behavioral dispositions. Narratives may therefore alter behavior by changing character, while embodied experiences may feed back into character and ultimately change the narrative self. The authors apply this framework to phenomena ranging from identification with fictional characters and stereotype priming to virtual-reality avatars and rehabilitation.

The article is interesting partly because it raises a broader question about how theoretical articles should be evaluated. Not all theories have the same epistemic status. In particular, it is useful to distinguish speculative theories from explanatory theories.

Speculative and explanatory theories

At one end of a continuum are speculative theories. These theories propose mechanisms or constructs that could potentially explain observations, but the empirical phenomena they are intended to explain have not themselves been firmly established. Freud’s psychoanalytic theories provide a familiar historical example. Concepts such as repression, the id, ego, and superego offered potentially interesting explanations of human behavior, but the empirical foundations for many of these explanations were weak and the theories were sufficiently flexible to accommodate many different observations.

At the other end are theories constructed to explain credible empirical findings. Perception research provides many examples. Psychophysical phenomena can often be demonstrated repeatedly under highly controlled conditions. Researchers may disagree about why a particular perceptual phenomenon occurs, but there is little disagreement that the phenomenon itself occurs. In this situation, theories compete to explain reliable data. The theoretical problem is not “Does this phenomenon really exist?” but “What mechanism explains it?”

This distinction can be represented on a five-point continuum:

  1. Highly speculative — the theory attempts to explain reported findings whose credibility or replicability has not been established.
  2. Mostly speculative — some credible empirical evidence exists, but important parts of the empirical foundation remain uncertain.
  3. Mixed — the theory integrates both well-established findings and substantially less certain findings.
  4. Mostly explanatory — the major empirical phenomena are supported by strong and reasonably replicable evidence.
  5. Explanatory theory of credible findings — the phenomena to be explained are firmly established; the principal uncertainty concerns their explanation rather than their existence.

An important implication of this distinction is that the number of empirical citations in a theoretical article does not necessarily make the theory empirically grounded. A theory can cite dozens of experiments and nevertheless remain highly speculative if those experiments are unreliable. What matters is the credibility of the phenomena that the theory attempts to explain.

Narrative embodiment

On this continuum, I would rate the narrative-embodiment framework 1 out of 5.

This does not mean that the theory is necessarily false. In fact, several aspects of it are intuitively plausible. It seems entirely possible that people’s narratives about themselves influence their behavior and that behavioral experiences subsequently influence how people think about themselves. The authors also make a constructive attempt to translate their framework into empirically measurable variables such as identification, self-efficacy, body image, and behavioral expression.

The problem is that the article does not critically evaluate the credibility of the empirical phenomena used to motivate the theory. Instead, published findings are typically treated as facts that require explanation.

The clearest example appears in the section on stereotype effects. The authors cite the famous study by Bargh, Chen, and Burrows (1996), in which participants exposed to words associated with old age reportedly walked more slowly after leaving the laboratory. Shimada et al. describe the result as follows:

“This demonstrates that stereotypes can shape behavior even without conscious awareness.”

The word “demonstrates” is important. The Bargh finding is not presented as a historically influential but controversial result. It is presented as empirical evidence for the proposed theoretical framework.

Yet the credibility of this finding has been questioned for well over a decade. Doyen et al. (2012) conducted a larger replication using automated measurement of walking speed. Their first experiment found essentially no difference between the elderly-prime and control conditions. The study included 120 participants, compared with 60 across Bargh et al.’s two original elderly-priming experiments, and used infrared sensors rather than manual stopwatch measurements. (PLOS)

Doyen et al.’s second experiment also showed that the outcome depended on experimenters’ expectations. The predicted slowing effect appeared when experimenters had been induced to expect primed participants to walk more slowly. Doyen et al. therefore concluded that priming alone was insufficient to reproduce the original walking-speed effect under their conditions. (PLOS)

Whatever one’s final judgment about behavioral priming, these findings clearly make the empirical status of the Bargh effect relevant to any theoretical review published in 2026. Yet Shimada et al. cite Bargh et al. without mentioning Doyen et al. or the broader controversy surrounding the replicability of behavioral priming.

This is not merely a minor omission in the reference list. It illustrates a fundamental problem in building psychological theories from the published literature. If significant findings are accepted at face value, virtually any collection of published results can provide apparent empirical support for a theoretical framework. After the replication crisis, however, the existence of a published finding cannot be equated with the existence of a credible phenomenon.

A more detailed examination of the history of the elderly-priming effect shows why this distinction matters. The evidence was much less consistent than the textbook version of the finding suggests, and later large preregistered studies provided additional negative evidence. A meta-analysis shows that the existing evidence is so weak and inconsistent that the true effect can range from no effect to a very strong effect without any known moderators (“Elderly Priming: Did It Ever Work?” ReplicabilityIndex)

A theory in search of credible phenomena

The same concern applies more broadly to the article. The authors draw on identification with fictional characters, stereotype effects, the Proteus effect, avatar embodiment, self-efficacy, narrative therapy, and other literatures. Some of these phenomena may ultimately prove more robust than others. Indeed, the authors cite meta-analytic evidence for the Proteus effect and occasionally acknowledge inconsistent findings. But there is no systematic attempt to distinguish highly credible phenomena from literatures characterized by small studies, selective reporting, or replication problems.

Consequently, the empirical literature functions primarily as illustration rather than as a set of established facts that constrain the theory.

This distinction is important because a genuinely explanatory theory should face constraints imposed by reliable observations. Suppose, for example, that repeated high-powered experiments established that identification with a fictional character reliably changed several character-consistent behaviors that had never been directly primed. A theory would then be needed to explain why these effects occur together. Shimada et al.’s concept of “character” might provide one possible explanation.

Indeed, one of the most promising ideas in the article points in precisely this direction. The authors suggest that activating one aspect of a character should affect other behaviors associated with the same character. Experiencing Superman’s ability to fly, for example, might subsequently increase helping behavior or bravery. If such cross-behavior effects were demonstrated reliably, especially under conditions that distinguished character identification from simpler priming or expectancy explanations, the narrative-embodiment framework would become substantially more explanatory and less speculative.

At present, however, the direction of inference is largely reversed. Published findings are assembled into a framework first, while the reliability of those findings is mostly taken for granted.

Conclusion

Narrative embodiment is an imaginative theoretical proposal. It integrates philosophical ideas about narrative identity and embodiment with several areas of psychological research and generates potentially testable hypotheses. For a journal called New Ideas in Psychology, this type of conceptual speculation is entirely appropriate.

But it is important to distinguish an interesting idea from an empirically established explanation.

On a continuum from speculative theory to explanation of credible empirical findings, I would rate the article:

1 / 5 — highly speculative.

Elderly Priming: Did It Ever Work?

The Short Version

Bargh, Chen, and Burrows (1996) reported that priming students with words related to old age made them walk more slowly, with effect sizes above one standard deviation (d = 1.04 and .79) across two studies of just 30 participants each. The finding became a cornerstone of behavioral-priming research and a famous casualty of the replication crisis after Doyen et al. (2012) failed to reproduce it. The Doyne et al. (2012) article came at the right time to cause a paradigm shift, but it was not the first replication failure.

The first replication failure appeared only a couple of years after the 1996 article: Dijksterhuis et al. (1998) found the same comparison at roughly a quarter of a standard deviation (d = .25 and .29), nonsignificant, and reinterpreted the disappearance as a side condition for a new contrast effect. Because social psychologists tracked the sign of an effect rather than its magnitude, a series of studies that did not reproduce Bargh’s result were published as successful demonstrations of moderators instead of as replication failures.

By 2012, elderly priming had been reported in 8 articles with 12 studies and 15 tests. I show that a meta-analysis at this time would have shown a wide prediction interval with possible effect sizes ranging from -0.1 to +1.0. Thus, Doyen et al.’s replication failure was entirely consistent with the full existing evidence. It was therefore entirely reasonable for Kahneman (2012) to ask for new and stronger evidence that Bargh never delivered. Since then, independent preregistered studies large failed to replicate past effects and priming theorists have largely abandoned priming research. The death of the priming paradigm provides a valuable lessons about the need to build theories on robust empirical foundations to avoid investing resources on phenomena that do not exist.

The Long Version

John A. Bargh studied with Robert Zajonc at the University of Michigan, earning his Ph.D. in 1981. Zajonc was an early champion of unconscious processes and used masked, subliminal presentations to study influences occurring outside conscious awareness. At New York University, Bargh developed a related program on automatic social cognition. At the time, studies from several laboratories suggested that stimuli presented outside awareness could influence feelings, judgments, and immediate reactions to other stimuli. Some of this evidence has since been challenged: meta-analyses find many subliminal effects hard to replicate, and concerns have been raised about publication bias and about whether participants were ever fully unaware of the masked stimuli.

Bargh matters for the history of psychology because his 1996 article made a much stronger claim. Rather than showing that unnoticed stimuli can shift immediate perceptions or evaluations, Bargh and colleagues claimed that activating a social concept could have a lasting influence on overt behavior without people being aware of that influence. In their famous elderly-priming experiment, students completed a scrambled-sentence task in which, in one condition, several words were related to old age and, in the control condition, were not. The words were visible, but participants were presumably unaware of the manipulation’s purpose or its possible effect on their behavior. Told to go to another room for a second study, participants then had their walking speed down the hallway measured as the real dependent variable. Bargh et al. reported that a few old-age words made students walk more slowly than controls.

The difference was not small. Effect sizes for two-group differences are often expressed in standard-deviation units. Cohen (1988) classified d = .50 as a medium effect and d = .80 as large. For comparison, an IQ test has a standard deviation of about 15 points, so half a standard deviation is 7.5 IQ points. In Bargh et al.’s first study, Experiment 2a, the effect was slightly larger than a full standard deviation, d = 1.04 — an exceptionally large effect for such a subtle manipulation.

Psychologists rarely run direct replications, but Bargh et al.’s article contained one. Experiment 2b used essentially the same procedure; again, primed participants walked significantly more slowly, and the effect stayed large, d = .79. The article thus appeared to offer unusually convincing evidence: two independent studies, same procedure, both producing significant and large effects.


The 1996 article spawned a large literature using primes such as money or God to influence behavior. Then, in 2012, Doyen et al. reported that they could not replicate the elderly-priming effect and suggested the original findings might be, at least partly, methodological artifacts. Walking speed in the original studies was timed manually with a stopwatch. And although the person timing walking speed was blind to condition, Doyen et al. noted it was unclear whether the experimenter who administered the priming task was also blind. If experimenters knew which participants had been primed, their expectations could have shaped participants’ behavior.

Doyen et al. tested this in a second experiment by manipulating experimenters’ expectations directly: some were led to expect primed participants to walk more slowly, others to walk faster. Objectively measured walking speed tracked those expectations — the elderly-prime effect appeared only when experimenters expected slowing — and the manual stopwatch measurements tracked them even more strongly. Doyen et al. thus provided experimental evidence that researchers’ expectations could produce the outcome of a behavioral-priming study.

Bargh responded in March 2012 with a Psychology Today post titled “Nothing in Their Heads.” He opened by noting that the 1996 finding had been theoretically predicted and fit a growing literature on automatic influences on behavior, but much of the post attacked the quality of Doyen et al.’s work and the peer review at PLOS ONE. Knowledgeable social-psychology editors and reviewers, he argued, would have caught the methodological problems; he had not been asked to review the paper, and implied that had he done so, he would have recommended rejection.

The tone drew wide criticism (Comments). Srivastava (2012), for one, characterized the episode as Bargh going “bananas.” The attack on the journal prompted a reply from the PLOS ONE editors on Bargh’s blog post (cf. Hodgkinson, 2012; Srivastava, 2012). Bargh later expressed regret over the tone and took the post down (Bartlett, 2013).

Bargh’s Defense of His Results in 2012

Bargh’s first substantive defense concerned experimenter effects. He stated that the experimenter in the original study had been blind to hypothesis and condition, and that a different person, also blind to condition, measured walking speed. If accurate, this substantially weakens Doyen et al.’s suggestion that experimenter expectations produced the original findings. But it does nothing for the more basic result: Doyen et al. used a larger sample and objective measurement and still failed to reproduce the effect.

Bargh therefore proposed several procedural differences to explain the replication failure. First, he argued that Doyen et al. had drawn participants’ attention to walking by telling them to “go straight down the hall when leaving,” and that making an automatic behavior conscious could eliminate the priming effect. But Doyen et al.’s article contains no such instruction; it says only that participants were “clearly directed to the end of the corridor” — and Bargh et al.’s own procedure had likewise directed participants toward the elevator down the hall. It is unclear the difference existed at all.

Second, Bargh questioned the strength of the manipulation: Doyen et al. put an elderly-related word in all 30 scrambled-sentence items, and Bargh argued that too many related words could make participants consciously aware of the theme and cancel an unconscious effect. Doyen et al. did find some evidence that participants could identify the elderly theme when directly probed — but nearly all denied noticing any connection between the sentence task and their walking.

Third, Bargh appealed to culture: priming works only if the association already exists in participants’ minds, so Belgian students might not link old age with slowness as American students do. Possible in principle — but Doyen et al. had adapted their materials to the Belgian sample by surveying 80 people about concepts associated with old age and selecting the frequent responses. Bargh et al.’s original article, by contrast, never measured whether their own participants associated the elderly stereotype with slower walking.

Finally, Bargh appealed to the accumulated literature: stereotype and behavioral priming had been replicated many times, so it was unreasonable to doubt the phenomenon over a single failure. This is persuasive only if the published record is an unbiased sample of all experiments run. That was precisely what the emerging replication crisis called into question. Journals favored novel, significant results and rarely published failed replications, so the sheer number of published successes could not reveal how often behavioral-priming experiments actually worked.

Doyen (2012) Was Not the First Failure

Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it, Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it,

It is a bit of a shame that many of the arguments Bargh uses in his criticism of the Doyen study are arbitrary, unsupported and, on occasion, false in light of other research (even some from the area of priming) (Nordbeck, 2012).

The earlier failures had been published, but not as failures. They appeared as successful demonstrations of moderators: conditions under which the effect was supposed to grow, shrink, or reverse (Cesario et al., 2006; Dijksterhuis et al., 1998; Hull et al., 2002). Bargh later cited these very studies without noting that none of them had reproduced his basic result — a difference between an elderly prime and a control group.

That this could happen reflects a habit of the field. Social psychologists tracked patterns of statistical significance more than the magnitude of effects. Evidence for a moderator requires a significant interaction, and once an interaction turns up, attention shifts from the main effect to the conditional effects that explain it. The claim is no longer “priming works” but “priming works differently under condition X” — and a study can support that claim while quietly failing to reproduce the main effect it was built on.

Dijksterhuis et al. (1998) is the clearest case, appearing just two years after the original. Studies 2a and 2b each contained the two conditions needed to test elderly priming — an elderly prime and a neutral prime, followed by the same judgment task. In both, primed participants walked slightly more slowly, but the effects were small and nonsignificant, d = .25 and .29, against Bargh’s d = 1.04 and .79. Same direction, a quarter of the magnitude.

But testing Bargh was not the point of their paper; behavioral contrast was. Some primed participants also judged a specific elderly exemplar — the Dutch Queen Mother, Princess Juliana — and then walked faster than both the neutral and the ordinary elderly-prime groups. That contrast effect was significant in both studies and became the headline. The authors did notice that their ordinary elderly prime had failed to reproduce Bargh’s assimilation effect, called it “somewhat surprising,” and proposed that the intervening judgment task had wiped it out — reading the small same-direction means as “residual assimilation.” The failure was not overlooked; it was reinterpreted into a footnote. By this reading, the first evidence that elderly priming is not as large or robust as advertised appeared fourteen years before Doyen.

Their explanation is itself revealing. If inserting an innocuous judgment task can cut an effect by roughly 75%, then the effect is extraordinarily sensitive to procedural detail. Researchers wanted theoretically predicted moderators that would specify when and why priming occurs. What the evidence kept pointing to instead were unknown moderators: incidental features of a procedure that swing the effect from large to nothing. A phenomenon that surfaces only under a narrow and poorly understood set of conditions is not the robust effect Bargh’s original experiments described.

Dijksterhuis was only the first. Hull et al. (2002) ran the paradigm in two studies and found priming only among participants high in self-consciousness; the overall effects were moderate but nonsignificant, ds = .57 and .54, in small samples. Cesario et al. (2006) added a youth prime that sped participants up, but the elderly-versus-control difference was small and nonsignificant, d = .23. Jeffries and Fazio (2008) found an interaction with a stopping rule on an anagram task and no main effect at all, d = −.05. N. Wyer (2011) produced the first robust replication — d = .70 on walking speed and d = 1.03 on working memory. And in 2012, the same year as Doyen, one article reported strong effects in one study (ds = 1.26, 1.05) but weak ones in another (ds = .29, .31).

Seen in this company, Doyen’s failure is unremarkable. Some studies produce large estimates and others produce weak estimates and all estimates in small samples have large sampling error and make estimation of the true effect size impossible. The common solution to imprecise estimates in small studies is to combine the studies in a meta-analysis.

A Meta-Analysis of Elderly Priming

Existing meta-analyses of priming pool many kinds of primes and outcomes, and the resulting heterogeneity tells us little about any one paradigm. So I conducted a meta-analysis restricted to elderly priming: 15 tests from 12 studies in 8 articles. Because publication bias is evident in the broader literature, I used a selection model that estimates the bias and returns a corrected effect size (Vevea & Hedges, 1995), with clustered bootstrapping to handle multiple tests nested within articles and to build the confidence and prediction intervals.

The bias-corrected effect was d = .42, 95% CI [−.01, .71]. The point estimate is close to the much larger meta-analysis of Dai et al. (2023), but the interval is wide and includes zero. The studies are small and individually imprecise, and the heterogeneity is itself barely pinned down: tau = .26, 95% CI [.00, .37]. Nothing here compels the conclusion that Bargh’s and Doyen’s results reflect different true effects.

The estimated mean and variation of the population effect sizes produce a 95% prediction interval that ranges form d = -0.10 to d = 1.03.

Thus, a replication study so large that sampling error is negligible could land anywhere from essentially nothing to a very large effect, because the existing evidence simply does not fix the size of the effect.

The meta-analysis also cannot explain how Bargh got two significant results from samples of N = 30. If the true effect were d = .42, a study that small has only about a 20% chance of reaching significance, so two independent significant results would occur about 4% of the time. Either Bargh was lucky, or something in his particular setup inflated the effect — for instance, a confound in the word list, if “Florida” slows NYU students’ walking by evoking a laid-back-Southern stereotype rather than an elderly one.

In short, a meta-analysis of the studies available in 2012 would have shown Bargh that Doyen’s result was entirely consistent with the broader evidence, however much it clashed with his own. There was no need to reach for incompetence. His deeper error was assuming that his two original studies, together with conceptual replications using other primes, had already established a general effect of elderly priming on behavior.

An important caveat is that publication-bias tests and corrections are difficult, especially when the set of studies is small. After Doyen et al.’s article was published, researchers shared anecdotal reports of additional unpublished replication failures (OSF Google Groups). If these unpublished studies could be recovered and included, they would likely reduce the average effect-size estimate.

Source: Comments to Ed Young’s Article
Source: Comments to Ed Young’s Article

The Fall-Out

Concerns about the credibility of social psychology were already building. Daniel Kahneman believed in priming and had featured it prominently in Thinking, Fast and Slow, yet he was growing worried about the field’s reputation and raised it with Bargh directly (Bartlett, 2012). In a widely circulated email to Bargh and other leading priming researchers, he proposed a simple remedy: if Doyen et al.’s failure was a fluke or a botch, the experts could settle the matter by running rigorous replications and showing the effects held. With an expected effect of d = .40, pooling resources for N = 400 would give roughly 98% power to detect it.

The collaboration never happened. Independent researchers ran preregistered replications instead, often with disappointing results — in Dai et al.’s (2023) meta-analysis, the preregistered studies, most of them replications, produced essentially no effect at all (d = .02).

In 2017, Kahneman returned to the problem, citing Overall (1969) on why underpowered research is not merely wasteful but pernicious: it inflates the share of false positives among published findings. Bargh’s experiments, with 15 participants per condition, had little power to detect any plausible effect and cannot account for the near-90% success rates in the published priming literature. Hundreds of studies manufactured the appearance of a robust phenomenon that low power, publication bias, heterogeneity, and near-zero preregistered effects do not support.

So the classic behavioral-priming paradigms simply faded from the research frontier. In that sense the field is practically dead — a stark case of strong belief outrunning weak evidence. Researchers can always find a reason to dismiss a failure and build ever more elaborate theory around inconsistent results, all without first establishing that the phenomenon the theory explains is real. As the meta-analysis above shows, even the 2012 evidence carried enough uncertainty to justify Kahneman’s call.

Which raises the obvious question: why wouldn’t a true believer, with the resources of a Yale professor, simply run the study and prove the point? Bartlett (2013) put it to Bargh directly:

So why not do an actual examination? Set up the same experiments again, with additional safeguards. It wouldn’t be terribly costly. No need for a grant to get undergraduates to unscramble sentences and stroll down a hallway.

Bargh was unenthusiastic. He wouldn’t ask graduate students worried about their job prospects to spend time on stigmatized research, and he was aware that some critics thought he had a “special touch” with priming — a word that sounds like praise but isn’t. “I don’t think anyone would believe me,” he said.

Source: Comments on Ed Young Article

The same article also reports an interview with Axel Cleeremans, one of Doyen’s collaborators. Cleeremans reports that his collaborators and he found found the tone of Bargh’s letter insulting, but that they also took his criticism seriously enough to conduct another replication study that addressed his stated concerns. The study also failed to replicate the original results. Another unpublished replication failure was posted online by Pashler et al. (2011; Srivastava, 2012)

While the evidence that elderly priming works was mixed and inconclusive, Bargh expressed strong confidence in his findings and his theory of unconscious activation of behavior. In his blog post, he believed that the long-term future of priming research was secure because the phenomenon rested on a large and converging empirical foundation:

Like many scientists, I take the long view and continue to have faith in science as a cumulative process, with particular studies that vary in their methods and approaches converging on deeper underlying principles and mechanisms. No single experiment, standing by itself, should be the basis for concluding anything, maybe most especially in psychological research. And when a single study does not replicate another one whose findings are solidly embedded in theories of more than one scientific field and which is consistent with dozens if not hundreds of other conceptual replications, then responsible scientists—and responsible science journalists—do not rush to judgment and make claims that the entire phenomenon in question is illusory. Thomas Kuhn famously argued in The Structure of Scientific Revolutions that it is the accumulation of contrary evidence over time, not single studies, which is needed to overturn established concepts and principles in any branch of science.

A decade on, his faith in cumulative science has been vindicated — just not as he expected. Failure after failure, and no convincing new evidence from the original researchers, has left claims about unconscious behavioral priming no better established than Freud’s speculations about the unconscious. Priming researchers published and circulated their successes and buried their failures, manufacturing the look of robustness; and as shown here, even the successes never told a clean story or gave clear evidence that elderly priming had ever produced a reliable effect.

Bargh was right about one thing: no single failure should have settled it. What mattered was the accumulation of evidence. That accumulation just happened to vindicate Doyen’s skepticism rather than his own confidence. Maybe, somewhere in his unconscious, Bargh knows this too — which might be why he never tried to replicate his 1996 findings.

The Lesson

Bargh is not the first scientist to spend much of a career on an idea that never lived up to its promise. That is an occupational hazard of working at the edge of a field: the questions are hard, the methods imperfect, and a genuine breakthrough is difficult to tell from a false lead until many technical problems are solved. Strong belief supplies the motivation to push through those obstacles — but it also makes the warning signs of failed studies easier to miss. Economists call inflated asset prices “irrational exuberance”: a stock can stay overvalued for as long as enough investors share the illusion, then crash to its true worth. Psychological theories kept aloft by shared belief rather than evidence behave the same way, as the citation history of Bargh’s 1996 article shows.

That casual attitude toward evidence is on display in Matthew Lieberman’s 2012 blog post on Doyen’s failure — a post that mentions, in passing, an unpublished non-replication from Lieberman’s own lab:

There have been multiple unpublished non-replications of the elderly-walking (including one in my lab that I may discuss in a future blog), but it is hard to know what to make of them given that they haven’t been peer-reviewed.

The catch-22 is clean. We trust only peer-reviewed findings, but psychology journals publish overwhelmingly significant results (Sterling, 1959; Sterling et al., 1995). So the failures stay invisible — unless they are dressed up as significant interactions, or land in a venue like PLOS ONE that isn’t gatekept by the same community. More telling still is what Lieberman took the stakes to be:

At a certain level it does not matter whether the exact primes Bargh used produce a change in walking speed over the exact distance he measured. Some have said ‘We need to replicate this exactly. Conceptual replications aren’t good enough’. But I’m not sure why we care about this specific manipulation unless we are about to start using it as an intervention to treat patients. What we care about is whether priming-induced automatic behavior in general is a real phenomenon. Does priming a concept verbally cause us to act as if we embody the concept within ourselves? The answer to this question is a resounding yes.

The “resounding yes” rests entirely on a published record that reports successes and buries failures — and, worse, treats interaction effects that masked the very replication failures at issue as further confirmation. Shown only the hits, it is easy to conclude the phenomenon must be real. Kahneman (2017) later admitted he had trusted that record too much when he built on it in his book. The lasting contribution of the controversy is that younger psychologists now treat direct replication as something more than an optional courtesy.

Younger scientists can lower the risk by studying the previous generation’s mistakes. We are fond of saying that each generation stands on the shoulders of giants. We say less often that many equally gifted scientists climbed just as high and backed ideas that proved wrong. The ones whose theories survived were not necessarily wiser or more careful at the outset; to a real extent, nature happened to cooperate. Darwin’s mechanism of natural selection withstood ever more stringent tests; Lamarck’s mechanism of inheritance did not. Neither man could have known which way it would go when he set out.

Because you cannot know in advance whether nature will cooperate, the main lesson of what Kahneman (2012) called the “train wreck” of priming research is methodological before it is anything else: do not expect clean answers from small samples, and do not trust significant results from journals that never publish the null ones. And beyond method, a matter of temperament — stay skeptical, and resist the pull to become a believer. As Feynman put it:

The first principle is that you must not fool yourself, and you are the easiest person to fool.