Blogging about statistical power, replicability, and the credibility of statistical results in psychology journals since 2014. Home of z-curve, a method to examine the credibility of published statistical results.
Show your support for open, independent, and trustworthy examination of psychological science by getting a free subscription. Register here.
“For generalization, psychologists must finally rely, as has been done in all the older sciences, on replication” (Cohen, 1994).
DEFINITION OF REPLICABILITY: In empirical studies with sampling error, replicability refers to the probability of a study with a significant result to produce a significant result again in an exact replication study of the first study using the same sample size and significance criterion (Schimmack, 2017).
See Reference List at the end for peer-reviewed publications.
Mission Statement
The purpose of the R-Index blog is to increase the replicability of published results in psychological science and to alert consumers of psychological research about problems in published articles.
To evaluate the credibility or “incredibility” of published research, my colleagues and I developed several statistical tools such as the Incredibility Test (Schimmack, 2012); the Test of Insufficient Variance (Schimmack, 2014), and z-curve (Version 1.0; Brunner & Schimmack, 2020; Version 2.0, Bartos & Schimmack, 2021).
I have used these tools to demonstrate that several claims in psychological articles are incredible (a.k.a., untrustworthy), starting with Bem’s (2011) outlandish claims of time-reversed causal pre-cognition (Schimmack, 2012). This article triggered a crisis of confidence in the credibility of psychology as a science.
Over the past decade it has become clear that many other seemingly robust findings are also highly questionable. For example, I showed that many claims in Nobel Laureate Daniel Kahneman’s book “Thinking: Fast and Slow” are based on shaky foundations (Schimmack, 2020). An entire book on unconscious priming effects, by John Bargh, also ignores replication failures and lacks credible evidence (Schimmack, 2017). The hypothesis that willpower is fueled by blood glucose and easily depleted is also not supported by empirical evidence (Schimmack, 2016). In general, many claims in social psychology are questionable and require new evidence to be considered scientific (Schimmack, 2020).
Each year I post new information about the replicability of research in 120 Psychology Journals (Schimmack, 2021). I also started providing information about the replicability of individual researchers and provide guidelines how to evaluate their published findings (Schimmack, 2021).
Replication is essential for an empirical science, but it is not sufficient. Psychology also has a validation crisis (Schimmack, 2021). That is, measures are often used before it has been demonstrate how well they measure something. For example, psychologists have claimed that they can measure individuals’ unconscious evaluations, but there is no evidence that unconscious evaluations even exist (Schimmack, 2021a, 2021b).
If you are interested in my story how I ended up becoming a meta-critic of psychological science, you can read it here (my journey).
References
Brunner, J., & Schimmack, U. (2020). Estimating population mean power under conditions of heterogeneity and selection for significance. Meta-Psychology, 4, MP.2018.874, 1-22 https://doi.org/10.15626/MP.2018.874
Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17, 551–566 http://dx.doi.org/10.1037/a0029487
Schimmack, U. (2020). A meta-psychological perspective on the decade of replication failures in social psychology. Canadian Psychology/Psychologie canadienne, 61(4), 364–376. https://doi.org/10.1037/cap0000246
Askarov, Z., Doucouliagos, A., Doucouliagos, H., & Stanley, T. D. (2024). Selective and (mis)leading economics journals: Meta-research evidence. Journal of Economic Surveys, 38(5), 1567–1592. https://doi.org/10.1111/joes.12598
Abstract
Askarov, Doucouliagos, Doucouliagos, and Stanley (2024) analyzed statistical power and excess statistical significance in a large collection of economics meta-analyses and concluded that much of the evidence reported in leading economics journals is potentially misleading. We used their open data to conduct a z-curve analysis to examine the credibility of economics using a different statistical model. Z-curve has several advantages over the power-analysis and Test of Excess Significance (TES) approach used by Askarov et al. First, it does not assume that all studies within a meta-analysis share a single population effect size. Instead, it models heterogeneity with a mixture model. Second, z-curve models selection for statistical significance and uses the fitted distribution of significant results to estimate the discovery rate that would be expected in the absence of selection.The discrepancy between the observed and expected discovery rates therefore provides a direct measure of selection bias. In contrast, TES does not explicitly model how selection distorts the distribution of observed effect sizes when estimating expected significance. Its UWLS estimator gives greater weight to more precise estimates, which typically come from larger samples. If smaller, less precise studies report inflated effect sizes, the weighted mean will be pulled toward the smaller effects observed in more precise studies, thereby reducing the estimated power assigned to the smaller studies. This weighting can reduce small-study bias, but it does not necessarily eliminate selection bias. Moreover, if true effect sizes systematically differ with study size, the same weighting can itself produce a biased estimate of the average effect. Third, z-curve distinguishes between overall power (the Expected Discovery Rate, EDR) and power conditioned on significance (the Expected Replication Rate, ERR). With heterogeneous data, the average power of significant results can be much higher than overall power. Finally, z-curve uses the EDR to obtain an upper bound on the false discovery rate using a formula developed by Sorić (1989).
First, Askarov et al.’s estimate-level mean power and the z-curve EDR are surprisingly similar, approximately 27% and 28%, respectively. A discovery rate of this magnitude implies a maximum false discovery rate of approximately 14%. Second, the expected replication rate of statistically significant results is approximately 70%, showing that the power of selected significant results is substantially higher than overall power. These estimates are similar to estimates obtained for randomized clinical trials in medicine and do not support pessimistic interpretations of this database based solely on its low median power. Low overall power is primarily a problem for discovery: true effects are less likely to reach significance, creating the potential for false negatives. Importantly, Askarov et al.’s own database shows that many nonsignificant estimates are nevertheless reported and incorporated into meta-analyses, where evidence can be aggregated to increase precision and statistical power. Thus, low power of individual studies does not by itself imply low credibility of the resulting literature.
Introduction
Concerns about the credibility of science are no longer purely academic. Scientific evidence informs consequential decisions about health, climate, and economic policy, making the credibility of published research important for both policymakers and the public. Yet academic incentives can undermine credibility. Researchers are rewarded for novel and statistically significant findings, whereas replications and corrections receive less attention. As a result, false positive findings may enter the literature and persist even when later evidence fails to support them.
Concerns about scientific credibility intensified after Ioannidis (2005) argued that most published research findings are false. Although influential, this claim was largely theoretical rather than based on an empirical estimate of false discoveries across science. For most significant results to be false positives, researchers must test many false hypotheses and have relatively low power to detect true effects. For example, if only 10% of tested hypotheses are true, statistical power is 50%, and the Type I error rate is 5%, then 5% of the true hypotheses and 4.5% of the false hypotheses will produce significant results. Consequently, nearly half of all significant results, 4.5/(4.5 + 5) = 47%, would be false discoveries.
Empirical investigations of scientific credibility have produced a less pessimistic but highly variable picture. Button et al. (2013) documented very low statistical power in neuroscience, with median power estimates across meta-analyses ranging from approximately 8% to 31%. In contrast, Jager and Leek (2014) analyzed reported p-values in major medical journals and estimated that only 14% of significant results were false discoveries. Direct replication projects introduced yet another measure of credibility. The Open Science Collaboration (2015) found that only 36% of psychology findings produced a significant result in the same direction in a replication, whereas Camerer et al. (2016) obtained a replication rate of 61% for laboratory experiments in economics. A much larger recent investigation of the social and behavioural sciences found that approximately half of tested claims replicated.
Concerns about credibility have also become prominent in economics. Large meta-research projects have documented selection for statistical significance and low statistical power. Most recently, Askarov et al. (2024) analyzed 368 meta-analyses containing 167,753 estimates, including 22,281 estimates published in 31 leading economics journals. They emphasized that median power in the leading journals was only 7% and reported substantial excess statistical significance, leading them to question the credibility of much published economics research. At the same time, direct replication and robustness studies have produced more encouraging results. Camerer et al. (2016) replicated 61% of experimental findings, while a recent large-scale study found that 72% of significant economics and political-science estimates remained significant and in the same direction under alternative analyses. Thus, empirical assessments of economics range from very low estimates of statistical power to substantially higher estimates of replicability and robustness.
These quantities can differ substantially when statistical power is heterogeneous. Moreover, estimates from different methods depend on different assumptions about effect-size heterogeneity, selection for significance, and the proportion of true null hypotheses. Consequently, apparently conflicting estimates of scientific credibility need not actually contradict one another.
The present study addresses this problem using z-curve, a statistical model that estimates several credibility parameters within a single coherent framework. Z-curve models heterogeneity in statistical power with a mixture distribution and explicitly models selection for statistical significance. Its main estimands are the EDR and the ERR. The EDR can be compared with the Observed Discovery Rate (ODR), the percentage of significant results, to assess and quantify selection for statistical significance. Furthermore, the EDR can be used to estimate the maximum False Discovery Risk (FDR) using a formula developed by Sorić (1989). We use the term risk rather than rate because the actual false discovery rate cannot be identified from the observed test statistics alone without knowing which tested null hypotheses are true.
Sorić’s formula shows that the relationship between EDR and maximum FDR is nonlinear. For example, an EDR of 20% implies a maximum FDR of approximately 21% at α=.05. Thus, even low mean discovery probabilities do not imply that most significant results are false positives.
Data
The Askarov et al. dataset combines 368 economics-related meta-analyses covering a broad range of research areas. The meta-analyses were identified through bibliographic databases, publisher websites, specialist journals, and searches of work by known meta-analysts; the search ended on July 31, 2021. When data were not publicly available, the authors contacted the original meta-analysts and obtained data from 74% of those contacted. To be included, a meta-analysis had to contain at least five primary studies and report both effect-size estimates and their standard errors. When multiple meta-analyses examined the same research area, the most recent and comprehensive one was selected. The final dataset contains 167,753 estimates, including 22,281 estimates published in 31 leading general-interest and field economics journals. The authors emphasize that the dataset is not necessarily representative of all empirical economics research, but rather of research areas that have been subjected to meta-analysis.
The database also contains identifiers for the original primary studies, making it possible to account for dependence among multiple estimates reported by the same study. The 167,753 estimates represent approximately 15,000 primary-study clusters.
Results
The most important estimate is the Expected Discovery Rate (EDR) of 27%. This estimate means that an unbiased sample of tests drawn from the same underlying population is expected to contain approximately 27% significant results. This estimate is surprisingly close to Askarov et al.’s estimate-level mean power of approximately 27%.
The two quantities are conceptually similar but not identical. Askarov et al. calculate directional power: significance is counted only in the direction of the estimated meta-analytic effect. Z-curve’s EDR uses two-sided statistical significance. Consequently, the null baseline for Askarov et al.’s directional calculation is 2.5%, rather than the conventional two-sided Type I error rate of 5%. The numerical difference between directional and two-sided power becomes very small as power increases, however, and does not explain the close agreement between the aggregate estimates.
The similarity of the mean estimates is particularly informative because the mean, rather than the median, determines the expected proportion of significant results.
For the full database, approximately 51% of reported estimates are significant, whereas z-curve estimates an EDR of 27%, a difference of approximately 24 percentage points. Askarov et al.’s estimate-level mean power for the full database is also approximately 27%, implying a very similar aggregate discrepancy between observed and expected significance. This numerical agreement should not be interpreted as validation of the two methods. Askarov et al. calculate power from a common meta-analytic effect within each research area, whereas z-curve estimates a heterogeneous distribution of noncentrality parameters. The two approaches can therefore produce very different results in individual heterogeneous meta-analyses even when their aggregate averages happen to agree.
It is unconventional to refer to the difference between observed and expected significance as a “rate of false positives.” The term false positive normally refers to a statistically significant result that incorrectly rejects a true null hypothesis. Excess significance does not establish that the excess results are false rejections of H0. They may instead reflect inflated estimates of real effects caused by selective reporting or specification searching. Thus, Askarov et al.’s excess-significance measure should not be interpreted as an estimate of the proportion of significant findings that are false discoveries.
In contrast, z-curve uses the EDR to estimate an upper bound on the proportion of significant results that could be false discoveries. Following Sorić (1989),FDRmax=(EDR1−1)1−αα.
With an EDR of 27%, the maximum FDR is approximately 14%. Allowing for sampling uncertainty in the EDR raises the upper confidence limit to approximately 19%. Thus, the results imply that no more than roughly one in five significant results could be false discoveries within the assumptions of the model. The actual FDR may be considerably lower. The Sorić bound is obtained under the extreme assumption that true alternatives are detected with perfect power; when power against true alternatives is lower, fewer of the observed significant findings can be attributed to true null hypotheses.
The most dramatic difference between Askarov et al.’s interpretation and the z-curve results concerns their emphasis on median power. Askarov et al. highlight median power of only 7% in leading economics journals and note that this value is close to the conventional 5% significance criterion. This comparison is misleading for two reasons.
First, their power calculation is directional. Under a true null hypothesis, their formula produces a probability of 2.5%, not 5%. Thus, a directional power estimate of 7% should not be compared directly with the two-sided Type I error rate of 5%. This distinction has little impact once power becomes moderate, but it matters for interpreting values very close to the null.
Second, and more importantly, median power is not the quantity that predicts how many significant results a literature should produce. The mean probability of significance does. Their own estimate-level mean power is approximately 27%, nearly four times their headline median of 7% and remarkably close to the z-curve EDR.
The distinction also matters for credibility. A low discovery probability across all tests implies that many results will be nonsignificant. This is a serious problem when nonsignificant findings are suppressed, because selective reporting will exaggerate the apparent success of the literature. But low discovery probability does not imply that significant findings themselves have similarly low replicability.
Z-curve estimates the Expected Replication Rate of significant results at 69%. Thus, although the EDR for all tests is only 27%, results that passed the significance threshold are estimated to have substantially higher power. The distinction follows directly from selection: results with higher underlying power are more likely to become significant and therefore are overrepresented among significant findings.
The ERR also includes any true null results that happened to become significant. At the maximum-FDR point estimate of 14%, the implied same-direction replication probability among the remaining true-positive results would be approximately 80%. This calculation should not be interpreted as a separate estimate of the true-positive power because the 14% FDR is itself an upper bound. It simply illustrates that low overall discovery probability can coexist with much higher replicability among significant results that reflect genuine effects.
In short, evaluations of credibility need to distinguish among several quantities: the probability of significance across all tests, the probability of significance among true alternatives, the replicability of results selected for significance, and the probability that a significant result is a false discovery. Median discovery probability provides little information about the latter two quantities.
Askarov et al.’s finding of low median power therefore does not by itself imply that economics research lacks credibility. Their own mean-power estimate and the z-curve EDR both suggest an underlying discovery probability of approximately 27%, while z-curve estimates an ERR of approximately 69% and a maximum FDR of approximately 14%. These results indicate substantial selection for statistical significance and considerable room for improvement, but they do not support the conclusion that the low median power of individual estimates, by itself, raises serious doubts about the credibility of the meta-analyzed economics literature.
Conclusion
In conclusion, meta-scientists often point out that extraordinary claims require extraordinary evidence and that academic incentives can reward researchers for making strong claims from weak evidence. Meta-science is not immune to these pressures. The claim that an entire discipline conducts studies with a typical probability of only 7% of rejecting a false null hypothesis is remarkable, if true. However, closer examination shows that this headline figure is a median discovery probability and is not the quantity that predicts the expected number of significant results or the credibility of significant findings. Askarov et al.’s own mean estimate is approximately 27%, closely matching the z-curve EDR, while z-curve estimates substantially higher replicability among significant results and a relatively modest upper bound on the false discovery rate. Thus, the evidence supports concerns about selective reporting and low discovery rates, but it does not support the much stronger conclusion that the low median power estimate by itself raises serious doubts about the credibility of economics research.
“Priming exercise” and psychological priming share a label, not necessarily a mechanism. Deliberate preparation for a known future competition provides no obvious need for unconscious goal activation, and the article presents no evidence that such activation actually occurs. The shared terminology leaves the proposed explanation primed for equivocation.
Holmberg and Kelly (2026) provide a useful critique of physiological explanations for “priming exercise,” but their attempt to connect this literature to psychological priming introduces a much less plausible mechanism. The connection appears to arise largely because the two literatures happen to use the same word. That shared terminology leaves the argument primed for equivocation.
In sports physiology, “priming exercise” refers to a deliberately performed bout of exercise intended to improve performance later that day. In cognitive and social psychology, priming refers to prior exposure to a stimulus that subsequently alters processing or behavior, sometimes without awareness or conscious intention. These are fundamentally different uses of the term. The fact that both involve something occurring before something else does not imply that they share a psychological mechanism.
This distinction becomes especially important when the authors invoke nonconscious goal priming. They suggest that exercise might activate performance goals and related behavioral representations that persist until later testing. They cite classic social-psychological priming research, including Bargh and colleagues, to support the possibility that goals can be activated without awareness and subsequently influence behavior.
But the proposed mechanism is poorly matched to the phenomenon being explained. An athlete does not ordinarily encounter a “priming exercise” incidentally. The athlete performs it because a competition or performance test is coming later. The later performance goal is therefore likely to have been activated before the exercise begins:
competition later → intention to prepare → priming exercise.
The goal is not plausibly dormant until the exercise somehow activates it unconsciously. Indeed, the goal is probably one of the reasons the athlete performs the exercise in the first place. During the exercise the athlete may also consciously think about the competition, technique, pacing, readiness, or expected benefits. Under those circumstances, invoking nonconscious goal activation is not merely unnecessary; the proposed causal sequence is almost backwards.
The authors’ own examples illustrate the problem. They discuss athletes rehearsing particular pacing strategies, using metronomes, receiving verbal cues, believing that squatting will improve subsequent jumping, and developing confidence or expectations about later performance. These are readily understood as deliberate preparation, task practice, expectancy, motivation, or attentional effects. None requires a nonconscious priming mechanism.
The scientific evidence offered for the nonconscious account is also weak. The article provides no exercise experiment demonstrating that a priming-exercise bout activates a previously inactive goal outside awareness and that this activation subsequently causes improved performance hours later. In fact, the authors repeatedly acknowledge that these possibilities “may” occur, are “hypothesized,” or “have yet to be directly examined.” Thus, the proposed mechanism is not an empirical finding from the exercise literature.
Instead, support is imported from a different literature on behavioral and goal priming. That literature is itself scientifically controversial, and citing classic demonstrations does not establish that the same mechanism operates in an entirely different situation involving intentional athletic preparation. The conceptual inference appears to be:
But identical terminology is not evidence of mechanistic equivalence.
The distinction matters because the paper already identifies much more plausible explanations. Task-specific practice, motor learning, expectancy, researcher effects, motivation, and explicit performance preparation could all produce later performance changes. These mechanisms fit the actual structure of the situation: athletes know that performance is coming and intentionally prepare for it. The nonconscious goal-priming hypothesis adds an unnecessary and poorly supported causal layer.
The problem can therefore be summarized simply: an implausible mechanism is invoked to explain effects that have not been shown to require that mechanism, and the empirical justification comes largely from a separate literature that happens to use the same word.
Mir, I. A. (2026). Influencer’s physical attractiveness and content aesthetics: Conscious and preconscious determinants of fashion-branded content engagement on Instagram. Journal of Creative Communications, 21(2), 201–219. https://doi.org/10.1177/09732586241288672
“The authors invoke an unproven perception–behavior mechanism to explain causation that was never observed. A direct path in a cross-sectional SEM establishes neither causation nor preconscious processing.”
Mir (2026) examines whether fashion influencers’ physical attractiveness and the aesthetics of their branded content predict followers’ engagement on Instagram. The study uses survey responses from 300 followers of 15 fashion influencers in Pakistan and analyzes the proposed relationships with structural equation modeling and mediation analyses. The main empirical finding is straightforward: followers who rate influencers and their content more positively also report more favorable attitudes and greater engagement. The methodological problem is that the article draws causal and psychological-process conclusions that the design cannot support.
The most serious problem is that all variables were measured in a single cross-sectional self-report survey. Participants simultaneously reported how attractive they considered the influencer, how aesthetically pleasing they considered the content, their attitude toward that content, and how often they viewed, liked, commented on, and shared it. Nothing was manipulated, and there was no temporal ordering of the variables. Nevertheless, the article repeatedly describes attractiveness and aesthetics as factors that “cause,” “trigger,” “stimulate,” or “activate” engagement. Those causal statements do not follow from the design.
For example, the proposed model assumes
attractiveness → attitude → engagement.
But the same covariance pattern is compatible with numerous alternatives. Followers who engage frequently with an influencer may develop more positive attitudes and subsequently rate that influencer as more attractive. A general liking or identification with the influencer could simultaneously increase attractiveness ratings, content-aesthetic ratings, attitudes, and engagement. The structural equation model cannot distinguish among these explanations.
The sampling procedure makes this problem particularly important. Participants had to have followed one of the selected influencers for more than six months. Thus, the sample is already conditioned on sustained interest in the influencer. People who disliked the influencer, found the content unattractive, or disengaged from it are systematically less likely to appear in the sample. The resulting correlations describe differences among an already selected group of followers; they provide weak evidence for the article’s practical recommendation that firms should hire physically attractive influencers because attractiveness causes engagement.
The article’s central methodological error is even more fundamental. It claims to distinguish a “conscious” route from a “preconscious” route. The mediated path
attractiveness/aesthetics → attitude → engagement
is interpreted as conscious influence, whereas a remaining direct path from attractiveness or aesthetics to engagement is interpreted as evidence of a preconscious perception–behavior process.
A direct regression coefficient is not a measure of unconscious processing.
If attractiveness predicts engagement after statistical adjustment for an explicit attitude measure, this merely shows residual covariance between those variables. That residual association could reflect measurement error in attitude, omitted mediators, stable preferences, common response tendencies, reverse causation, or numerous other processes. Nothing in the study measures awareness, intention, automaticity, processing speed, or participants’ ability to report the causes of their behavior. Consequently, the data provide no evidence that engagement was “preconscious” or unintentional.
For the same reason, the indirect path through an explicit attitude measure does not establish a conscious causal mechanism. Participants consciously completed the attitude questionnaire, but that does not mean the psychological process producing their engagement operated consciously. Statistical mediation and conscious psychological mediation are different concepts.
This problem is especially consequential because the “preconscious” interpretation is one of the article’s principal theoretical contributions. The authors explicitly invoke the perception–behavior literature, including Bargh et al. (1996), to justify the claim that a significant direct path demonstrates automatic behavior. But the current research contains none of the experimental procedures that would be required to test an automatic perception–behavior effect. The SEM therefore cannot adjudicate between conscious and unconscious processes.
The measures themselves also create substantial interpretive problems. On page 209, physical attractiveness is measured with “stylish,” “good looking,” “sexy,” and “elegant.” Content aesthetics is measured with “striking,” “wonderful,” “fascinating,” and “lovely,” while attitude is measured with “pleasant,” “good,” “likeable,” and “my favourite.” These constructs are conceptually and evaluatively intertwined. “Wonderful” and “lovely,” for example, are not narrowly aesthetic judgments, while “stylish” and “elegant” are not purely measures of physical attractiveness. Much of the model may therefore reflect a broad positive-evaluation factor rather than distinct psychological constructs connected by causal pathways.
The observed correlations are consistent with this concern. Physical attractiveness correlates .64 with content aesthetics and .64 with attitude; content aesthetics correlates .66 with attitude and .68 with reported content consumption. Demonstrating discriminant validity with the Fornell–Larcker criterion does not eliminate the possibility that halo effects or general liking strongly influence all of these ratings.
The attempt to dismiss common-method bias is also inadequate. All predictors, mediator variables, and outcomes were obtained from the same respondent at the same time using similar rating formats. The authors test whether several sets of items can be represented by a single latent factor and conclude that poor single-factor fit shows that common-method bias is not important. That conclusion does not follow. Common-method variance does not require every item to load on one factor. Several distinguishable constructs can coexist while correlations among them are inflated by shared method, evaluative consistency, acquiescence, or halo effects.
The proposed “snowball effect” suffers from the same causal problem. The authors find that self-reported consumption behaviors—viewing, reading comments, and liking—predict contribution behaviors such as commenting and sharing, and conclude that consumption gradually causes users to progress toward more active participation. Yet consumption and contribution were measured simultaneously. Someone who frequently comments and shares influencer content almost necessarily also views and consumes that content. A positive cross-sectional association therefore does not demonstrate a temporal progression from passive to active engagement. Testing a snowball process would require longitudinal evidence showing that earlier consumption predicts subsequent increases in contribution.
There may also be an unmodeled dependence problem. The 300 respondents followed one of only 15 macro-influencers. Followers of the same influencer are not necessarily independent observations. Influencers may differ systematically in appearance, production quality, follower demographics, posting frequency, and baseline engagement. Those influencer-level characteristics could generate correlations among respondent ratings. The reported SEM appears to treat all 300 followers as independent rather than accounting for clustering by influencer.
Another reporting issue concerns the engagement scale. The response categories are described as 1 = “very often,” 2 = “often,” 3 = “sometimes,” and 4 = “never.” Thus, larger numerical values indicate less engagement. Yet positive path coefficients are consistently interpreted as greater attractiveness and aesthetics producing greater engagement. The article does not clearly state in the reported method that these scores were reverse-coded. If they were reversed before analysis, that transformation should have been explicitly documented. If they were not, the substantive interpretation of the coefficients would be reversed.
Finally, the statistical success of the model should not be confused with strong evidence for the hypotheses. All five proposed hypotheses are supported, including the weakest direct attractiveness effect, β = .12, t = 2.19. The study was not preregistered, and the sample-size justification consists largely of the statement that N = 300 is sufficient for purposive sampling and structural equation modeling rather than an a priori power analysis tied to the focal effects. Good model-fit indices demonstrate that a specified covariance model can reproduce the observed covariance matrix; they do not establish that the arrows in Figure 2 represent the true causal processes.
The study therefore supports a much narrower conclusion than the article claims. Among existing long-term followers of fashion influencers, positive ratings of influencer attractiveness and content aesthetics are associated with positive attitudes and greater self-reported engagement. That descriptive association is plausible and potentially useful.
The study does not establish that physical attractiveness or content aesthetics cause engagement, that attitude mediates these effects causally, that any residual direct relationship reflects a preconscious perception–behavior mechanism, or that passive engagement develops over time into active contribution.
A suitable experimental design would manipulate influencer attractiveness and content aesthetics independently, randomly assign participants to conditions, measure actual engagement behavior, and include measures capable of testing awareness or automaticity. A longitudinal design would be required to test the proposed consumption-to-contribution “snowball” process. Without such evidence, the article’s strongest psychological claims are interpretations imposed on cross-sectional correlations rather than findings produced by the research design.
Costa, T. (2026). The Bayesian audit: Evaluating the proportionality of scientific claims to evidence—a case study on social priming and walking speed. Frontiers in Psychology, 17, 1799078. DOI: 10.3389/fpsyg.2026.1799078
This article applies Bayes’ theorem to one conveniently selected t value and calls the result a ‘Bayesian audit.’ Most readers may simply ignore it because it appeared in Frontiers in Psychology. Those who want a more substantive reason can point to this review.
Costa (2026) introduces a “Bayesian audit,” a six-step framework intended to evaluate whether the strength of scientific claims is proportional to the evidence supporting them. The idea is sensible. Statistical significance does not tell us how strongly we should believe a scientific claim, and surprising claims based on weak evidence deserve particularly careful scrutiny. Costa illustrates the proposed method with one of social psychology’s most famous findings: Bargh, Chen, and Burrows’s (1996) claim that priming college students with words related to old age caused them to walk more slowly afterward.
Unfortunately, the audit itself is problematic. It misrepresents important features of the original study, considers only a fraction of the available evidence, and reduces a question about the magnitude and robustness of an effect to a comparison between a null and an inadequately specified alternative hypothesis.
The first problem is surprisingly basic. Costa describes the original finding as based on a study with approximately t(28) = 2.0 and p ≈ .05. But Bargh et al. actually reported two elderly-priming experiments. In Experiment 2a, the comparison was t(28) = 2.86, p < .01. They then conducted Experiment 2b as a replication and again reported slower walking, t(28) = 2.16, p < .05. Costa appears to approximate the weaker second result while failing to mention the stronger first result or even that the original article contained two studies.
That is an odd starting point for an audit. If the purpose is to reconstruct how much evidence supported the claim in 1996, both original studies should be included.
Costa also incorrectly describes participants as being “subliminally exposed to words related to old age.” They were not. Participants consciously read words while completing a scrambled-sentence task. The claimed unconscious component was that participants supposedly did not realize that the elderly-related words subsequently affected their walking. Bargh et al. themselves explicitly distinguished this procedure from subliminal priming; Experiment 3 of their paper used genuinely subliminal presentation of faces.
This distinction matters because Costa uses the apparent implausibility of unconscious effects on motor behavior to motivate skeptical prior probabilities. One should at least characterize the causal claim correctly before assigning a prior to it.
The treatment of replication evidence is even more problematic.
Costa cites Doyen et al. (2012) and Harris et al. (2013) as subsequent replication attempts. Doyen et al. did replicate the elderly-walking paradigm. In a substantially larger study using automated measurement, they found essentially no priming effect. Their second experiment further suggested that experimenter expectations could influence the result.
Harris et al. (2013), however, did not replicate elderly priming at all. They attempted to replicate Bargh et al.’s 2001 high-performance goal-priming experiments, in which achievement words were supposed to improve performance on a cognitive task. Calling Harris et al. a replication of the elderly-walking finding is simply an error.
More importantly, why is a Bayesian audit conducted in 2026 based primarily on one t statistic from 1996?
There is now a substantial literature on behavioral priming. Dai et al. (2023), for example, meta-analyzed 351 studies and 862 effect sizes and concluded that behavioral priming effects could be detected across a large literature. Conversely, Mac Giolla et al. (2024) examined 70 close replication attempts of 49 social-priming findings. Ninety-four percent produced smaller effects than the originals, only 17% were significant in the predicted direction, and among 52 replications conducted without an original author, none was significant in the original direction; the pooled effect for those independent replications was essentially zero.
These sources do not necessarily settle the question. Meta-analyses themselves can be distorted by publication bias and other forms of selection. But that is precisely why an audit should examine them critically. An audit of a 30-year-old scientific claim should evaluate the accumulated evidence, not simply convert one selected original result into a Bayes factor.
There is an even more fundamental problem with the statistical question Costa asks.
Costa assigns prior probabilities of .05, .10, and .20 to the alternative hypothesis and combines these with an estimated Bayes factor of approximately 3. This yields posterior probabilities of .14, .25, and .43, respectively. The arithmetic is straightforward. The interpretation is not.
Why should the prior probability that the effect exists be .05 or .10?
Costa acknowledges that these values are illustrative rather than derived from an elicitation procedure. But these priors largely determine the conclusion that posterior belief remains low. Starting with a 5% probability and multiplying the prior odds by a Bayes factor of 3 inevitably produces a low posterior probability.
More importantly, what exactly is the hypothesis whose prior probability is 5%?
There is a major difference between these propositions:
elderly-related words have exactly zero effect on walking speed;
elderly-related words have some nonzero effect;
elderly-related words have a psychologically meaningful effect;
elderly priming produces effects of the magnitude originally reported;
The scientifically interesting issue today is probably not whether the population effect is exactly zero. The effect could be d = .05 or d = .10. Such an effect would make the point null hypothesis technically false while providing little support for the dramatic theoretical interpretation of the original experiments.
This is why effect sizes matter. Bargh et al.’s original studies implied very large effects. Subsequent evidence raises the possibility that the true effect, if it exists at all, is much smaller. A useful audit therefore needs to ask how large the effect is and how precisely it has been estimated—not merely whether H0 or H1 receives the larger Bayes factor.
Costa’s procedure also conflates two different kinds of priors. One is the prior model probability: how likely H1 is relative to H0 before seeing the data. The other is the prior distribution over possible effect sizes within H1. A Bayes factor for a composite alternative necessarily depends on the latter. Yet the article emphasizes sensitivity to prior model probabilities while giving much less attention to the effect-size assumptions used to obtain BF₁₀ ≈ 3.
This creates another problem with Costa’s distinction between “evidence” and “belief.” He describes the Bayes factor as quantifying evidence supplied by the data and posterior probability as combining this evidence with prior belief. But a Bayes factor is not simply a property of the observed data. It depends on the statistical models being compared, including the distribution of effect sizes assumed under the alternative hypothesis.
There is also an internal inconsistency in the treatment of the replication evidence. Costa states that the later replication attempts yielded Bayes factors close to 1 and therefore had little evidential impact. That makes sense: a Bayes factor of 1 leaves prior odds unchanged. Yet the subsequent synthesis says that posterior belief “collapses under replication.” It cannot do both. Replications with BF ≈ 1 cannot cause posterior belief to collapse. To demonstrate such a decline, one would need Bayes factors favoring the null or another competing model and then accumulate this evidence formally.
Publication bias is another conspicuous omission. The article is motivated by the replication crisis and explicitly acknowledges that biased data limit the usefulness of evidential measures. Yet the actual Bayesian calculation treats the published Bargh result as though it were an observation selected independently of statistical significance.
That is unrealistic. A BF of 3 obtained from a randomly selected study and a BF of 3 obtained from a literature in which statistically significant and theoretically exciting findings were preferentially published do not have the same evidential implications. If selection contributed to the replication crisis, an audit of the original published evidence needs to take selection seriously.
Costa also describes the original study’s low statistical power as an additional reason for skepticism. Low power certainly matters because significant results from low-powered studies tend to exaggerate effect sizes, particularly in a selected literature. But sample size has already entered the likelihood used to compute the Bayes factor. Low power is therefore not independent evidence against the hypothesis. The additional concern arises from selection, analytic flexibility, measurement error, and effect-size inflation.
The deeper problem is that Costa reduces a scientific question to H0 versus H1 when several competing explanations exist. Doyen et al.’s work raised experimenter expectancy as one possible explanation. Other possibilities include a genuinely small priming effect, effects restricted to particular conditions or individuals, procedural artifacts, or some combination of these mechanisms. A Bayes factor contrasting an exact-zero model with a generic nonzero-effect model cannot determine which causal explanation is correct.
This is especially important because rejecting H0 would not establish Bargh’s theory. Even convincing evidence for a tiny difference in walking speed would not demonstrate that automatic stereotype activation generally controls overt behavior.
The Bayesian audit is therefore based on a reasonable principle but a poor demonstration. Scientific claims should indeed be proportional to evidence. The problem is that assessing proportionality requires accurately identifying the original evidence, considering the accumulated replication literature, evaluating publication bias, distinguishing statistical from substantive hypotheses, and estimating plausible effect sizes and their uncertainty.
Ironically, the elderly-priming case illustrates the weakness of Costa’s audit more effectively than it illustrates its strengths. A proper audit should not ask merely whether one selected t statistic changes the odds that an effect is exactly zero. It should ask what three decades of evidence tell us about the magnitude, robustness, boundary conditions, and causal interpretation of the phenomenon.
On those questions, uncertainty remains. There may be a small elderly-priming effect. The evidence does not establish that the effect is exactly zero. But neither does the accumulated evidence support taking the spectacular effects reported in 1996 at face value. The important scientific task is to estimate what effect remains after accounting for uncertainty and bias. That requires more than Bayes’ theorem applied to one conveniently chosen t value.
Przybylinski, E. (2026). Whatever you say: Changing transference-based problem behavior with if–then plans. Self and Identity. Advance online publication. https://doi.org/10.1080/15298868.2026.2613846
Przybylinski (2026) reports two experiments examining whether implementation intentions can prevent problematic behaviors triggered by transference. The theoretical logic is straightforward: subtle resemblance to a significant other is assumed to activate that person’s representation automatically, which can then influence memory, goals, and behavior. An if–then plan is proposed to prevent the activated representation from guiding subsequent behavior.
The principal concern is the credibility of the evidence on which this argument rests.
The article treats automatic behavioral priming as a well-established foundation. For example, it cites Bargh et al. (1996), Bargh et al. (2001), Chartrand and Bargh (1999), and related studies as evidence that contextual cues can automatically trigger overt behavior without awareness or intention. Yet behavioral priming is precisely one of the areas most affected by the replication crisis. The article does not discuss this change in evidential status. Thus, evidence that was considered persuasive when these studies were conducted is largely presented in 2026 as though subsequent replication failures had not occurred.
This matters because the two experiments are themselves products of that earlier research era. The author explicitly states that they were conducted as dissertation research in the late 2000s, before preregistration became standard. Study 1 included only 60 participants, or 20 participants in each of the three strategy conditions, despite testing interaction hypotheses. Study 2 included 47 participants. The manuscript provides a power justification, but because the studies were not preregistered, it is unclear whether the reported power analysis reflects a prospectively specified design decision. The reported “post hoc power” provides little additional information.
The results are remarkably successful. The focal interactions involving behavioral readiness, memory, and overt submissive behavior are all statistically significant in the predicted direction. The reported effects are also very large, frequently exceeding d = 1 and reaching d = 1.69 in Study 1. Across the major focal tests, the success rate is effectively 100%, while average observed power based on the reported effects is roughly 80%.
A perfect success rate is not impossible when power is 80%, but it is more successful than expected. Schimmack (2012) emphasized that unusually high success rates relative to estimated power can indicate that published effect sizes and success rates should not be taken at face value. Here, the outcomes within each experiment are dependent, so a simple excess-success or incredibility calculation would not be appropriate. With only two studies, there is no statistical smoking gun. Nevertheless, the combination of small samples, large effects, multiple significant focal outcomes, and absence of preregistration warrants substantial caution.
The historical timing makes this concern more important. These experiments were conducted approximately 15 years before their publication. During that interval, psychology experienced a replication crisis that directly challenged the credibility of the behavioral-priming literature on which the article relies. Yet the 2026 article does not supplement the old experiments with a contemporary, adequately powered, preregistered replication.
This is particularly striking because the central experiment is readily replicable. The author remains at the same institution where the original research was conducted, and the paradigm requires undergraduate participants rather than an unusually difficult population. A preregistered replication with a substantially larger sample could have provided highly informative evidence about whether the large effects observed in the original dissertation studies survive contemporary scrutiny.
The absence of such a replication changes how the evidence should be interpreted. The results are not invalid merely because they were collected before the replication crisis. But neither the reported effect sizes nor the perfect pattern of statistical success should be treated as reliable estimates of the underlying effects without independent replication.
Klein, J. W., & Swann, W. B., Jr. (2026). Social psychology’s empty-self metaphor and the replication crisis. Perspectives on Psychological Science, 21(2), 138–153. https://doi.org/10.1177/17456916251401849.
Klein and Swann (2026) offer an interesting diagnosis of the replication crisis, but the evidence does not support the strength of their theoretical interpretation.
Their central empirical observation is striking. They coded 41 hypotheses from Many Labs 1 and 2 as either consistent with an “empty-self” metaphor or not. None of the nine “empty-self” hypotheses replicated, whereas 26 of 32 “not-empty-self” hypotheses replicated, a difference of 81 percentage points. Given the small number of empty-self studies, however, the estimate is much less precise than the point estimate suggests; an approximate 95% confidence interval for the difference is about 47 to 91 percentage points. The authors appropriately describe the result as preliminary.
The more serious problem is interpretation. The analysis is correlational. Studies were not randomly assigned to use an “empty-self” theory, and the authors did not code plausible confounding variables that could themselves predict replication success. These include the sample size and statistical power of the original study, the strength of the situational manipulation, whether the manipulation was consciously perceived, the causal proximity between manipulation and outcome, and the prior plausibility of the predicted effect. Their own coding examples illustrate the problem. A gray background affecting support for austerity is coded as “empty self,” whereas paying for a workshop affecting attendance is coded as “not empty self.” The latter is still a situational effect. What differs most obviously is that payment is a strong and behaviorally relevant manipulation, whereas background color is a weak and remote one.
Thus, the empirical result may show that studies proposing large effects of weak, incidental situational manipulations replicate poorly. That is interesting, but it is not the same as showing that studies fail because they neglect an enduring self.
The theoretical explanation is also speculative and at times internally strained. Klein and Swann sometimes treat replication failures as evidence that subtle situational manipulations have little or no effect. In the Many Labs studies, this inference can be justified when very large samples produce narrow confidence intervals around zero. But this cannot be generalized to all failed replications. In other literatures, including elderly priming, the available confidence intervals may still be compatible with small effects. Failure to obtain significance is not equivalent to demonstrating an effect of exactly zero.
At the same time, Klein and Swann suggest that subtle situational effects may depend on the person. They argue that people may respond only to cues to which they are “tuned” and that primes may work only when they connect to existing self-representations. That possibility is not new to priming theory. Priming effects have long been assumed to depend on whether participants possess the relevant stereotype or representation, and some priming studies have explicitly predicted interactions—for example, effects of religious primes that depend on participants’ religiosity.
But this moderation account has an important implication. If a prime affects people for whom it is relevant and has little effect on others, the population-average effect should normally be reduced, not eliminated. Unless one assumes theoretically unusual crossover interactions in which the prime produces effects in opposite directions for different people, sufficiently large studies should still detect a small average effect. For elderly priming, for example, it is easy to imagine that some participants might be more responsive to an elderly stereotype than others. It is much harder to explain why the same prime should make another substantial group walk faster.
This creates a useful empirical question that Klein and Swann do not examine. Their 41 hypotheses should be coded not only as “empty self” or “not empty self,” but also according to whether the original hypothesis predicted a main effect or a Person × Situation interaction. If nearly all of the “empty-self” studies in Many Labs tested simple main effects, that matters because the larger priming literature already contains many moderator and interaction hypotheses. The Many Labs sample would then not represent the full theoretical range of priming research. More broadly, the authors should distinguish studies proposing universal effects from studies explicitly predicting conditional effects.
A stronger analysis would therefore code each study independently for situation strength, awareness of the manipulation, personal relevance, causal distance between manipulation and outcome, original sample size and power, prior plausibility, and whether the prediction was a main effect or an interaction. Only then could one determine whether an “empty-self” construct predicts replication after plausible alternative explanations have been taken into account.
The irony is that Klein and Swann criticize social psychology for building theories on weak evidence, but their own “empty-self” explanation is itself a speculative theory supported by weakly diagnostic data. The empirical pattern—0% versus 81% replication—is interesting. The claim that this pattern is caused by neglect of an enduring self is not established.
On a 1-to-5 scale from speculative theory to theory explaining highly credible phenomena, I would rate the article about 2/5. The phenomenon to be explained is credible: some classes of social-psychological findings replicate poorly. The proposed explanation—that they fail because social psychology adopted an “empty-self” metaphor—remains largely speculative.
“Theories are a dime a dozen” (Ed Diener, personal communication)
Bingley, W. J., Worthy, P., Wiles, J., & Haslam, S. A. (2026). A social identity theory of digital identity. Perspectives on Psychological Science, 21(4), 346–382. https://doi.org/10.1177/17456916261419813.
Bingley and colleagues (2026) propose a new “social digital identity theory” (SDIT) to explain how social identities operate across online, offline, and hybrid environments. The theory is unusually explicit: the authors formulate 20 propositions concerning identity salience, digital platforms, embodiment, well-being, group functioning, polarization, culture, and other outcomes.
This is not a judgment that the theory is uninteresting or implausible. A rating of 2 means that much of the distinctive theory remains speculative and that the empirical findings used to motivate it have not been systematically evaluated for credibility.
There are good features. Bingley et al. clearly distinguish established ideas from novel hypotheses. They explicitly acknowledge that many of their propositions—particularly P4 through P11, P18, and P20—have not previously been tested and require empirical evaluation. Other propositions are extensions of the much older social-identity and self-categorization literature. Thus, they do not claim that all 20 propositions are established facts.
The problem is different. When published studies are cited as empirical support, the authors rarely ask how strong that evidence actually is. There is no systematic consideration of statistical power, independent replication, publication bias, or whether an apparently supportive literature consists mainly of selected significant findings.
Embodiment provides a good example
Proposition 5 states that social-identity salience is shaped by embodiment. To motivate this proposition, the authors cite studies suggesting that heat activates anger concepts, social rejection makes rooms feel colder, keeping secrets feels physically burdensome, fist clenching activates related concepts, and facial-muscle activation influences experience. They then extend this reasoning to virtual embodiment and the Proteus effect.
A particularly revealing example concerns elderly priming.
The authors cite a virtual-reality study by Reinhard et al. in which participants embodied an older or younger avatar. Participants who had embodied the older avatar subsequently walked more slowly during the first part of a walking test than participants who embodied the younger avatar, p = .033. However, these differences disappeared during the second half of the walk.
Bingley et al. describe this result as operating similarly to “classic priming studies in social psychology” and cite Bargh, Chen, and Burrows (1996).
That citation is problematic because Bargh et al.’s famous claim that elderly-related primes cause people to walk more slowly became one of the best-known examples of the replication problems in social psychology. Doyen et al. (2012) conducted a larger replication using automated measurement and failed to reproduce the original effect in their first experiment.
None of this replication history is mentioned.
The result is an evidential chain that looks stronger on paper than it really is:
reported embodiment effects → elderly-avatar study → classic elderly priming → support for an embodiment proposition.
There are several citations, but citations are not the same thing as strong evidence.
The elderly-avatar study may be interesting, and the possibility that virtual embodiment affects subsequent behavior certainly deserves further study. But a marginally significant result that is then linked to a famous but poorly replicated priming finding does not provide strong evidence for a broad theoretical proposition about embodiment and social identity.
The same caution applies to meta-analytic evidence. A meta-analysis can summarize a literature very precisely while still giving a misleading estimate if the underlying literature is affected by publication bias and selective reporting. The important question is therefore not simply whether a theory can cite studies—or even a meta-analysis—that support a proposition. The question is whether the phenomenon itself has been demonstrated with credible evidence.
Why 2 rather than 1?
SDIT deserves more than the lowest score because it is not free-form speculation. It builds on established theoretical traditions, makes explicit predictions, distinguishes some novel claims from older ones, and repeatedly identifies propositions that require future empirical tests.
But it does not deserve a middle or high score because much of the distinctive theory is not yet explaining firmly established empirical phenomena. Moreover, the embodiment section demonstrates that the authors sometimes treat published findings as evidence without critically assessing whether those findings survived the replication crisis.
Thus, the 2/5 rating means:
SDIT is a structured and testable theoretical proposal with some grounding in established research, but much of its distinctive content remains speculative, and the empirical literature used to motivate some propositions is treated too uncritically.
The citation of Bargh et al. (1996) is a particularly clear example. In 2026, a theory article should not invoke elderly priming as supportive evidence without informing readers that the original phenomenon itself remains empirically uncertain.
Aytürk, E., & Saribay, S. A. (2026). Neglect of the individual as a neglected problem: The relevance of combined idiographic-nomothetic approaches for social psychology today. Self and Identity. https://doi.org/10.1080/15298868.2026.2697941
Aytürk and Saribay (2026) make a useful methodological argument in “Neglect of the individual as a neglected problem.” Social psychology typically averages across people even though the same situation may have different psychological meanings for different individuals. They advocate greater use of personalized stimuli and intensive repeated-measures designs to identify person-specific processes before generalizing across people.
They also suggest that this problem might help explain replication failures in social priming. For example, “elderly” might evoke a frail grandparent for one person and a vigorous retiree for another. Averaging across these individuals could obscure different or even opposing effects. This is a reasonable hypothesis. It is also useful that the authors explicitly acknowledge low statistical power, questionable research practices, and publication bias as established contributors to replication failures.
The problem is that they nevertheless cite Bargh et al. (1996) rather uncritically. The famous finding that activating the elderly stereotype makes people walk more slowly is treated as an example of a potentially heterogeneous priming effect, without informing readers that the original effect became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original effect in their first experiment.
There were earlier attempts to explain the inconsistent findings by invoking individual differences. For example, Cesario et al. (2006) reported that elderly priming slowed participants with relatively positive attitudes toward elderly people but sped up those with negative attitudes. Other studies similarly reported moderator effects. But these small-sample studies do not provide strong evidence that a robust main effect was merely hidden by heterogeneity. Cesario et al.’s study, for example, had only about 67 participants for the relevant analysis and did not obtain a conventionally significant overall priming effect. The evidential burden therefore shifted to an interaction estimated from an even less powerful design.
This is important because interaction effects are generally more difficult to estimate precisely than main effects. Finding a significant personality moderator in a small study does not establish that a failed main effect was really caused by individual differences. The moderator itself needs adequate power and independent replication. A more detailed discussion of these purported replications of elderly priming is available in “Elderly Priming: Did It Ever Work?”
Here Aytürk and Saribay’s methodological proposal may actually provide a better way forward. Intensive repeated measurement can obtain much more information from each participant and therefore can study within-person Person × Situation effects with fewer participants than a conventional design may require. This comes at a cost: many observations per person are needed, and ecological or experience-sampling studies typically sacrifice some of the experimental control available in tightly controlled laboratory experiments. Moreover, any between-person moderator still ultimately depends on having enough individuals.
Thus, the proposal itself is worth pursuing. Low power means that the failed priming literature often provides lack of evidence rather than definitive evidence of absence. Small, conditional, person-specific priming effects remain possible.
But they remain hypotheses to be demonstrated. Individual heterogeneity can explain variation in a real effect; it cannot by itself establish that an unreliable effect is real. A 2026 article explicitly concerned with the replication crisis should therefore not cite Bargh et al. (1996) without acknowledging the replication failures that make the existence of the phenomenon itself uncertain.
Smolinski, J., Smolinski, R., Kesting, P., & Kröcher, F. (2026). Ethical guidelines for designing, developing, and deploying AI negotiation agents. Group Decision and Negotiation, 35, 56. https://doi.org/10.1007/s10726-026-10011-2
One of the article’s central psychological concerns is that AI negotiation agents could exploit human vulnerabilities through sublimiminal or unconscious influence. This possibility sounds alarming because it suggests that an AI agent might manipulate a negotiator without the person even realizing that an influence attempt is taking place.
But the psychological evidence cited for this possibility is surprisingly weak.
The authors cite Bargh et al. (1996), one of the best-known studies of unconscious behavioral priming. Its famous elderly-priming experiments reported that participants exposed to words associated with old age subsequently walked more slowly. Yet the article does not mention that this finding later became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original priming effect in their first experiment. The subsequent history of the finding raises further questions about whether elderly priming was ever a reliable phenomenon; I review that evidence in more detail in “Elderly Priming: Did It Ever Work?”
The problem is broader than the Bargh study. Subliminal persuasion has never been shown to be a particularly powerful method for changing consequential consumer choices. An older meta-analysis of subliminal advertising estimated an effect of only r = .059 on consumer choice. Later positive studies suggest that effects depend on restrictive boundary conditions, but even these results are not robust. Thus, it remains scientifically unknown if or when subliminal attempts to manipulate people actually work.
This raises an important question for the ethical analysis. Why would an AI negotiation agent need subliminal manipulation at all?
Ordinary persuasion operates in full awareness and can be far more direct. A seller can explicitly frame an offer as prestigious, scarce, safe, innovative, or socially desirable. A brand can persuade consumers that “Apple is cool,” and a consumer may knowingly pay considerably more for an Apple product because that identity has value to them. Nothing has to be flashed below the threshold of awareness. The person can see the message, understand the message, and still be influenced by it.
That form of influence is potentially much more relevant to AI negotiation agents. An AI can learn what arguments appeal to a particular person, emphasize some information rather than other information, frame alternatives strategically, appeal to identity or social norms, create urgency, flatter, build trust, or exploit known preferences. None of these processes requires a controversial theory of unconscious behavioral priming.
This distinction matters because the paper arguably directs attention toward the more sensational but less empirically credible risk. The image of an AI system exploiting subliminal psychological mechanisms is disturbing, but current psychology provides little evidence that such techniques constitute a powerful means of controlling behavior. More ordinary forms of conscious persuasion are both more plausible and potentially more consequential.
The ethical concern is therefore legitimate, but the psychological rationale should be reformulated. Rather than claiming that AI will be able to exploit “unconscious priming far more precisely” than humans, a more defensible concern is that AI may become exceptionally effective at personalized persuasion using psychological processes that operate with the target’s awareness.
The problem with citing Bargh is that it directs attention to irrational fears about subliminal manipulation, when AI may use much more powerful strategies to manipulate people with information that is in plain sight.
Shimada, S., Tanaka, S., Morioka, S., Rode, G., Roy, J.-M., & Rossetti, Y. (2026). Narrative embodiment: A conceptual framework linking the narrative self and the embodied self. New Ideas in Psychology, 83, 101285. https://doi.org/10.1016/j.newideapsych.2026.101285
Shimada and colleagues (2026) propose “narrative embodiment” as a conceptual framework linking two aspects of the self: the narrative self and the embodied self. Their central proposal is that “character,” borrowed from Paul Ricoeur, functions as an intermediary between narratives about who we are and relatively stable behavioral dispositions. Narratives may therefore alter behavior by changing character, while embodied experiences may feed back into character and ultimately change the narrative self. The authors apply this framework to phenomena ranging from identification with fictional characters and stereotype priming to virtual-reality avatars and rehabilitation.
The article is interesting partly because it raises a broader question about how theoretical articles should be evaluated. Not all theories have the same epistemic status. In particular, it is useful to distinguish speculative theories from explanatory theories.
Speculative and explanatory theories
At one end of a continuum are speculative theories. These theories propose mechanisms or constructs that could potentially explain observations, but the empirical phenomena they are intended to explain have not themselves been firmly established. Freud’s psychoanalytic theories provide a familiar historical example. Concepts such as repression, the id, ego, and superego offered potentially interesting explanations of human behavior, but the empirical foundations for many of these explanations were weak and the theories were sufficiently flexible to accommodate many different observations.
At the other end are theories constructed to explain credible empirical findings. Perception research provides many examples. Psychophysical phenomena can often be demonstrated repeatedly under highly controlled conditions. Researchers may disagree about why a particular perceptual phenomenon occurs, but there is little disagreement that the phenomenon itself occurs. In this situation, theories compete to explain reliable data. The theoretical problem is not “Does this phenomenon really exist?” but “What mechanism explains it?”
This distinction can be represented on a five-point continuum:
Highly speculative — the theory attempts to explain reported findings whose credibility or replicability has not been established.
Mostly speculative — some credible empirical evidence exists, but important parts of the empirical foundation remain uncertain.
Mixed — the theory integrates both well-established findings and substantially less certain findings.
Mostly explanatory — the major empirical phenomena are supported by strong and reasonably replicable evidence.
Explanatory theory of credible findings — the phenomena to be explained are firmly established; the principal uncertainty concerns their explanation rather than their existence.
An important implication of this distinction is that the number of empirical citations in a theoretical article does not necessarily make the theory empirically grounded. A theory can cite dozens of experiments and nevertheless remain highly speculative if those experiments are unreliable. What matters is the credibility of the phenomena that the theory attempts to explain.
Narrative embodiment
On this continuum, I would rate the narrative-embodiment framework 1 out of 5.
This does not mean that the theory is necessarily false. In fact, several aspects of it are intuitively plausible. It seems entirely possible that people’s narratives about themselves influence their behavior and that behavioral experiences subsequently influence how people think about themselves. The authors also make a constructive attempt to translate their framework into empirically measurable variables such as identification, self-efficacy, body image, and behavioral expression.
The problem is that the article does not critically evaluate the credibility of the empirical phenomena used to motivate the theory. Instead, published findings are typically treated as facts that require explanation.
The clearest example appears in the section on stereotype effects. The authors cite the famous study by Bargh, Chen, and Burrows (1996), in which participants exposed to words associated with old age reportedly walked more slowly after leaving the laboratory. Shimada et al. describe the result as follows:
“This demonstrates that stereotypes can shape behavior even without conscious awareness.”
The word “demonstrates” is important. The Bargh finding is not presented as a historically influential but controversial result. It is presented as empirical evidence for the proposed theoretical framework.
Yet the credibility of this finding has been questioned for well over a decade. Doyen et al. (2012) conducted a larger replication using automated measurement of walking speed. Their first experiment found essentially no difference between the elderly-prime and control conditions. The study included 120 participants, compared with 60 across Bargh et al.’s two original elderly-priming experiments, and used infrared sensors rather than manual stopwatch measurements. (PLOS)
Doyen et al.’s second experiment also showed that the outcome depended on experimenters’ expectations. The predicted slowing effect appeared when experimenters had been induced to expect primed participants to walk more slowly. Doyen et al. therefore concluded that priming alone was insufficient to reproduce the original walking-speed effect under their conditions. (PLOS)
Whatever one’s final judgment about behavioral priming, these findings clearly make the empirical status of the Bargh effect relevant to any theoretical review published in 2026. Yet Shimada et al. cite Bargh et al. without mentioning Doyen et al. or the broader controversy surrounding the replicability of behavioral priming.
This is not merely a minor omission in the reference list. It illustrates a fundamental problem in building psychological theories from the published literature. If significant findings are accepted at face value, virtually any collection of published results can provide apparent empirical support for a theoretical framework. After the replication crisis, however, the existence of a published finding cannot be equated with the existence of a credible phenomenon.
A more detailed examination of the history of the elderly-priming effect shows why this distinction matters. The evidence was much less consistent than the textbook version of the finding suggests, and later large preregistered studies provided additional negative evidence. A meta-analysis shows that the existing evidence is so weak and inconsistent that the true effect can range from no effect to a very strong effect without any known moderators (“Elderly Priming: Did It Ever Work?” ReplicabilityIndex)
A theory in search of credible phenomena
The same concern applies more broadly to the article. The authors draw on identification with fictional characters, stereotype effects, the Proteus effect, avatar embodiment, self-efficacy, narrative therapy, and other literatures. Some of these phenomena may ultimately prove more robust than others. Indeed, the authors cite meta-analytic evidence for the Proteus effect and occasionally acknowledge inconsistent findings. But there is no systematic attempt to distinguish highly credible phenomena from literatures characterized by small studies, selective reporting, or replication problems.
Consequently, the empirical literature functions primarily as illustration rather than as a set of established facts that constrain the theory.
This distinction is important because a genuinely explanatory theory should face constraints imposed by reliable observations. Suppose, for example, that repeated high-powered experiments established that identification with a fictional character reliably changed several character-consistent behaviors that had never been directly primed. A theory would then be needed to explain why these effects occur together. Shimada et al.’s concept of “character” might provide one possible explanation.
Indeed, one of the most promising ideas in the article points in precisely this direction. The authors suggest that activating one aspect of a character should affect other behaviors associated with the same character. Experiencing Superman’s ability to fly, for example, might subsequently increase helping behavior or bravery. If such cross-behavior effects were demonstrated reliably, especially under conditions that distinguished character identification from simpler priming or expectancy explanations, the narrative-embodiment framework would become substantially more explanatory and less speculative.
At present, however, the direction of inference is largely reversed. Published findings are assembled into a framework first, while the reliability of those findings is mostly taken for granted.
Conclusion
Narrative embodiment is an imaginative theoretical proposal. It integrates philosophical ideas about narrative identity and embodiment with several areas of psychological research and generates potentially testable hypotheses. For a journal called New Ideas in Psychology, this type of conceptual speculation is entirely appropriate.
But it is important to distinguish an interesting idea from an empirically established explanation.
On a continuum from speculative theory to explanation of credible empirical findings, I would rate the article:
1 / 5 — highly speculative.
Cookie Consent
We use cookies to improve your experience on our site. By using our site, you consent to cookies.