Klein, J. W., & Swann, W. B., Jr. (2026). Social psychology’s empty-self metaphor and the replication crisis. Perspectives on Psychological Science, 21(2), 138–153. https://doi.org/10.1177/17456916251401849.
Klein and Swann (2026) offer an interesting diagnosis of the replication crisis, but the evidence does not support the strength of their theoretical interpretation.
Their central empirical observation is striking. They coded 41 hypotheses from Many Labs 1 and 2 as either consistent with an “empty-self” metaphor or not. None of the nine “empty-self” hypotheses replicated, whereas 26 of 32 “not-empty-self” hypotheses replicated, a difference of 81 percentage points. Given the small number of empty-self studies, however, the estimate is much less precise than the point estimate suggests; an approximate 95% confidence interval for the difference is about 47 to 91 percentage points. The authors appropriately describe the result as preliminary.
The more serious problem is interpretation. The analysis is correlational. Studies were not randomly assigned to use an “empty-self” theory, and the authors did not code plausible confounding variables that could themselves predict replication success. These include the sample size and statistical power of the original study, the strength of the situational manipulation, whether the manipulation was consciously perceived, the causal proximity between manipulation and outcome, and the prior plausibility of the predicted effect. Their own coding examples illustrate the problem. A gray background affecting support for austerity is coded as “empty self,” whereas paying for a workshop affecting attendance is coded as “not empty self.” The latter is still a situational effect. What differs most obviously is that payment is a strong and behaviorally relevant manipulation, whereas background color is a weak and remote one.
Thus, the empirical result may show that studies proposing large effects of weak, incidental situational manipulations replicate poorly. That is interesting, but it is not the same as showing that studies fail because they neglect an enduring self.
The theoretical explanation is also speculative and at times internally strained. Klein and Swann sometimes treat replication failures as evidence that subtle situational manipulations have little or no effect. In the Many Labs studies, this inference can be justified when very large samples produce narrow confidence intervals around zero. But this cannot be generalized to all failed replications. In other literatures, including elderly priming, the available confidence intervals may still be compatible with small effects. Failure to obtain significance is not equivalent to demonstrating an effect of exactly zero.
At the same time, Klein and Swann suggest that subtle situational effects may depend on the person. They argue that people may respond only to cues to which they are “tuned” and that primes may work only when they connect to existing self-representations. That possibility is not new to priming theory. Priming effects have long been assumed to depend on whether participants possess the relevant stereotype or representation, and some priming studies have explicitly predicted interactions—for example, effects of religious primes that depend on participants’ religiosity.
But this moderation account has an important implication. If a prime affects people for whom it is relevant and has little effect on others, the population-average effect should normally be reduced, not eliminated. Unless one assumes theoretically unusual crossover interactions in which the prime produces effects in opposite directions for different people, sufficiently large studies should still detect a small average effect. For elderly priming, for example, it is easy to imagine that some participants might be more responsive to an elderly stereotype than others. It is much harder to explain why the same prime should make another substantial group walk faster.
This creates a useful empirical question that Klein and Swann do not examine. Their 41 hypotheses should be coded not only as “empty self” or “not empty self,” but also according to whether the original hypothesis predicted a main effect or a Person × Situation interaction. If nearly all of the “empty-self” studies in Many Labs tested simple main effects, that matters because the larger priming literature already contains many moderator and interaction hypotheses. The Many Labs sample would then not represent the full theoretical range of priming research. More broadly, the authors should distinguish studies proposing universal effects from studies explicitly predicting conditional effects.
A stronger analysis would therefore code each study independently for situation strength, awareness of the manipulation, personal relevance, causal distance between manipulation and outcome, original sample size and power, prior plausibility, and whether the prediction was a main effect or an interaction. Only then could one determine whether an “empty-self” construct predicts replication after plausible alternative explanations have been taken into account.
The irony is that Klein and Swann criticize social psychology for building theories on weak evidence, but their own “empty-self” explanation is itself a speculative theory supported by weakly diagnostic data. The empirical pattern—0% versus 81% replication—is interesting. The claim that this pattern is caused by neglect of an enduring self is not established.
On a 1-to-5 scale from speculative theory to theory explaining highly credible phenomena, I would rate the article about 2/5. The phenomenon to be explained is credible: some classes of social-psychological findings replicate poorly. The proposed explanation—that they fail because social psychology adopted an “empty-self” metaphor—remains largely speculative.