Category Archives: Uncategorized

How Effective is Psychotherapy for the Treatment of Depression?

Conclusion: A meta-analysis of clinical trials in Western nations that compare psychotherapy to treatment as usual shows an average effect size of half a standard deviation with 95% of effects ranging from a quarter to three-quarters of a standard deviation. Thus, psychotherapy is an important component of treatment for depression.

Introduction

Psychotherapy works, p<.05p < .05.

However, a statistically significant result alone does not really help practitioners and patients assess the benefits of psychotherapy. The important question is not “Is the effect greater than zero?” but “How much does psychotherapy help?”

A single study cannot answer this question precisely because psychotherapy studies tend to have relatively small samples and therefore substantial sampling error. Meta-analysis was developed to address this problem. A simple meta-analysis combines the effect-size estimates from individual studies while taking their sampling error (standard errors) into account to obtain a more precise estimate of the average effect. As the number of studies increases, sampling error in the average estimate can become practically negligible.

The headline estimate of psychotherapy effectiveness is about 70% of a standard deviation on a measure of clinical depression such as the Beck Depression Inventory (g=.72g=.72), with very little sampling error (SE=.03SE=.03; Plessen et al., 2023). This suggests that psychotherapy has a substantial positive effect. However, although the average effect can be estimated very precisely, two other sources of uncertainty undermine the usefulness of this estimate: (a) publication bias and (b) variation in true effect sizes across populations, treatments, control conditions, and other study characteristics.

Publication Bias

One problem is publication bias. Studies that show that psychotherapy is effective may be more likely to be published than studies that fail to show an effect. If so, the published literature will exaggerate effectiveness. Plessen et al. (2023) therefore included PET–PEESE, a statistical method designed to estimate the effect after accounting for a possible relationship between effect size and sampling error.

The method exploits the fact that small studies need larger estimated effects to achieve statistical significance, whereas large studies can achieve significance with smaller estimated effects. If statistically significant findings are preferentially published, this creates a relationship between sampling error and observed effect sizes. PET–PEESE uses this relationship to estimate what the effect would be as sampling error approaches zero. Across the PET–PEESE analyses in Plessen et al.’s multiverse, the estimated effect averaged only g=.18g=.18.

The problem is that publication bias is not the only reason why effect sizes might be related to study size. Small and large studies may differ systematically in other ways. For example, smaller studies might provide more intensive and costly treatments, whereas larger trials might use briefer or online interventions. Control groups, patient populations, and other study characteristics may also differ with study size. If these characteristics genuinely influence treatment effects, PET–PEESE can mistake real differences in treatment effectiveness for publication bias and adjust the effect downward too strongly.

We are therefore left with remarkably different answers to a seemingly simple question. A conventional meta-analysis suggests an effect of g=.72g=.72, whereas PET–PEESE suggests an effect closer to g=.18g=.18. Is psychotherapy highly effective, only modestly effective, or somewhere in between?

Fortunately, other methods can help distinguish publication bias from genuine variation in treatment effects. The first aim of this blog post is to use these methods to obtain a more credible estimate of the effectiveness of psychotherapy.

Heterogeneity

The second problem is heterogeneity. The effectiveness of psychotherapy may vary across populations, types of treatment, control conditions, and other study characteristics. In other words, there may be no single effect size that describes the effectiveness of psychotherapy under all conditions. An average can still be calculated, but it may not provide a useful prediction for any particular treatment setting.

In meta-analysis, variation in the true effect sizes across studies is called heterogeneity. In Plessen et al.’s full three-level meta-analysis, the average effect was g = .72, but the between-study variance was tau² = .364, corresponding to tau = .60. Thus, although the average effect was estimated very precisely, the true effects varied substantially from study to study.

Assuming a normal distribution of true effects, approximately 95% of study-level effects would be expected to fall between g = -.46 and g = 1.90. Thus, the same meta-analysis that estimates the average effect very precisely also allows for true effects ranging from moderately favoring the control condition to extremely large benefits of psychotherapy.

This wide range shows why a precisely estimated average does not necessarily answer the question, “How effective is psychotherapy?” An average of g = .72 tells us that psychotherapy is beneficial on average, but it provides little guidance about the effect we should expect in a particular population, treatment, or comparison condition.

To obtain more informative estimates, we need to understand why treatment effects vary across studies. The second aim of this blog post is therefore to identify one or more groups of studies with reasonably similar true effect sizes and to estimate how effective psychotherapy is under more specific conditions.

This goal may seem counterintuitive because a general rule in statistics is that larger samples are more informative. This is clearly true within an individual study: larger samples reduce random sampling error and produce more precise estimates. In meta-analysis, however, simply adding more studies does not necessarily make the answer more informative. Adding studies reduces sampling error in the average effect, but it can also increase heterogeneity if the added studies examine different populations, treatments, countries, or control conditions.

In this sense, less can be more. A meta-analysis of a smaller but more comparable set of studies may provide a more useful estimate than a much larger meta-analysis that averages over systematically different conditions. For example, an estimate of the effectiveness of cognitive behavioral therapy for adults within a particular healthcare setting and relative to a particular control condition may be more informative than a single average that combines different countries, therapies, populations, and comparison groups.

The goal is therefore not to make the meta-analysis as large as possible, but to define groups of studies for which an average effect has a clear and useful interpretation.

Munder et al.’s Meta-Analysis

The Study

I used Munder et al.’s (2022) meta-analysis because it examined a particularly plausible moderator: the treatment received by patients in the control group. In psychotherapy trials, the control condition plays a role similar to the comparison condition in a drug trial. Some patients are assigned to a waitlist and may receive little or no treatment during the study. In other trials, patients in the control group receive treatment as usual (TAU), and psychotherapy is added only for the treatment group. For example, both groups may receive antidepressant medication, while only the treatment group also receives psychotherapy.

Treatment as usual can also vary substantially in intensity and effectiveness. The more effective the treatment received by the control group, the smaller the additional benefit of psychotherapy is likely to be. This means that an average effect size that combines waitlist controls with different forms of treatment as usual may obscure meaningful differences in effectiveness.

Another plausible moderator is the country in which the study was conducted. Even without a specific hypothesis about cultural differences in the treatment of depression, psychological effects often vary across countries and cultures. Country can therefore serve as a useful proxy for differences in culture, healthcare systems, recruitment practices, and other contextual factors that may influence treatment effects.

A third possible moderator is the type of psychotherapy. Although meta-analyses suggest that many forms of psychotherapy are effective, it remains possible that some treatments produce larger effects than others.

The goal is not only to determine whether these moderators predict effect sizes. Even more important is to determine how much heterogeneity remains after accounting for them. If control condition, country, and treatment type explain a substantial portion of the variation across studies, we can obtain more informative estimates for specific sets of conditions—for example, the expected effect of CBT in U.S. studies compared with a waitlist control group.

Random Effects Meta-Analysis

I first analyzed the data with a standard random-effects meta-analysis using the R package metafor. This model estimates the average effect size and the amount of heterogeneity while assuming that the observed studies are not distorted by publication bias. The analysis produced an average effect of g = .72, SE = .08, together with substantial heterogeneity, tau = .67. These results closely replicate the findings of Plessen et al. (2023): psychotherapy appears highly effective on average, but treatment effects vary greatly across studies.

PET

I next applied PET, one of the regression-based methods used by Plessen et al. (2023) to address publication bias. PET tests whether effect-size estimates are related to their standard errors and extrapolates this relationship to a hypothetical study with no sampling error.

The analysis showed a strong relationship between effect sizes and sampling error, b = 3.29, SE = .57. The estimated intercept at zero sampling error was slightly negative, g = -.07, SE = .15, and not significantly different from zero. Taken literally, PET would therefore suggest that there is no convincing evidence for an average psychotherapy effect after correcting for publication bias.

However, this interpretation depends critically on the assumption that the relationship between effect size and sampling error is caused by publication bias. Moreover, substantial heterogeneity remained even after fitting PET, tau = .57. Thus, the model still allows for large positive effects in some studies while simultaneously estimating an average effect close to zero.

zcurve3

To examine publication bias with fewer assumptions about the relationship between effect size and study size, I also analyzed the data with zcurve3. Z-curve converts each effect size and its standard error into a z-value. A two-sided z-value greater than 1.96 is statistically significant. Z-curve uses the distribution of statistical evidence, particularly the significant z-values, to estimate the underlying distribution of evidential strength and to predict how many significant and non-significant results should be observed.

Figure 1 shows that the fitted distribution predicts somewhat more non-significant results than were actually observed. The expected discovery rate—the proportion of statistically significant findings predicted by the model—is only 39%, whereas the observed discovery rate is higher. This pattern is consistent with some selection for statistical significance. However, unlike PET, z-curve is very uncertain about the magnitude of this bias. The 95% confidence interval for the expected discovery rate extends as high as 83%, so the data are also compatible with little or no excess of significant findings.

Z-curve also reveals substantial variation in the strength of evidence across studies. Studies with non-significant z-values have low estimated power, whereas many of the statistically significant studies have moderate to high power, ranging from approximately 62% to 98%. The average estimated power of the significant studies—the Expected Replication Rate—is 74%. Z-curve also estimates the maximum proportion of statistically significant findings that could be false positives. Although the upper bound of the 95% confidence interval reaches 42%, even this conservative bound implies that the majority of statistically significant findings are unlikely to be false positives.

Zcurve3 can also provide estimates on the effect-size scale. The estimated overall mean effect is g = .33, with a wide 95% confidence interval ranging from g = .12 to g = .90. This interval contains both the small PET estimate and substantially larger conventional random-effects estimates. Rather than forcing the data toward one of these answers, zcurve3 makes the uncertainty about publication bias explicit, rather than assuming that bias is large (PET) or that there is no bias (RMA).

Making Sense of Heterogeneity

The z-curve plot shows heterogeneity in the strength of evidence below the x-axis. These values are estimates of local statistical power for studies with z-values in the corresponding ranges. Studies with small z-values have low estimated power and therefore provide little information about the magnitude of the underlying true effect. In the present data, local power is below 50% throughout the non-significant range. Around z = 2, however, estimated local power rises above 50%, reaching 62% for results just above the conventional significance thresho ld and increasing further for larger z-values.

This provides a principled way to distinguish relatively informative from highly uncertain effect-size estimates. Importantly, the criterion is not statistical significance itself. The criterion is estimated local power. In these data, the point at which local power exceeds 50% happens to coincide approximately with the conventional significance threshold. Thus, focusing on the statistically significant results in this particular dataset amounts to focusing on the subset for which zcurve3 estimates that there is more signal than noise.

zcurve3 also estimates the maximum false-positive rate. The upper bound of the 95% confidence interval is 39%. Thus, even under a conservative interpretation, the majority of results in this more informative subset are estimated to reflect a genuine positive treatment effect.

The main advantage of zcurve3 is that it is designed to examine heterogeneity. There are two types of heterogeneity to consider. Heterogeneity in the strength of evidence and heterogeneity in effect sizes.

The z-curve plot shows heterogeneity in the strength of evidence below the x-axis. These values are local power estimates for the corresponding ranges of z-values. Average power is low for non-significant results. These are mostly studies with small samples and large sampling error, and they provide little information about the magnitude of the underlying true effects. In contrast, studies with z-values greater than about 2 have considerably greater evidential strength. For results just above the conventional significance threshold, estimated local power is already 62% and increases further for larger z-values.

zcurve3 also estimates the maximum false-positive rate. The upper limit of the 95% confidence interval is 39%. Thus, even under this conservative estimate, the majority of the significant results are expected to reflect a genuine positive treatment effect. It is therefore possible to identify a subset of studies that provides substantially stronger evidence about treatment effectiveness.

zcurve3 also provides effect-size estimates for subsets of studies defined by their observed z-values. For the statistically significant results, the estimated mean effect size is g = 1.41, but uncertainty is substantial, 95% CI [.56, 2.21]. Heterogeneity among these effects is also very large, tau = .98, with a wide 95% CI [.21, 1.67]. Thus, focusing on studies with stronger evidence does not solve the heterogeneity problem. We still need to ask why some studies produce much larger effects than others. This requires examining potential moderators—that is, study characteristics that explain variation in effect sizes—with the goal of identifying one or more groups of studies for which an average effect size has a meaningful interpretation.

It therefore makes sense to examine studies with strong evidence more closely. A complication is that unusually large observed effects may partly reflect sampling error or selection. zcurve3 addresses this problem by using empirical-Bayes shrinkage to produce adjusted effect-size estimates that pull unusually large observed effects toward more plausible underlying values. Figure 2 shows these adjusted estimates.

The first five effects, representing four studies, still appear unusually large even after this adjustment. These studies may be scientifically interesting and deserve careful examination and replication, but they are not necessarily informative about the average effect in the broader group of studies. Instead, if their unusually large effects arise from study characteristics that are not shared by the remaining studies, combining them with the rest simply increases unexplained heterogeneity.

Chiang et al. provides a useful example. It is the only study from Taiwan in this meta-analysis. Consequently, its exceptionally large effect is perfectly confounded with the individual study: with only one Taiwanese study, we cannot determine whether the effect reflects Taiwan, some other feature of the study, or sampling variation. The study is informative about that particular Taiwanese sample, but it cannot establish a general Taiwanese treatment effect and it contributes little to estimating the effect for a population of Western studies. For that purpose, including it mainly adds heterogeneity that cannot be explained or generalized.

To produce a more homogeneous and interpretable set of studies, I applied several additional restrictions. First, I removed five studies with unique characteristics or questionable reporting that made their unusually large effects difficult to interpret or generalize. Second, I removed studies with large sampling error (SE > .30). This is actually a relatively modest restriction compared with Stanley, Jarrell, and Doucouliagos (2010), who proposed estimating meta-analytic effects from only the most precise 10% of published estimates when publication selection is a concern. Their argument is that highly imprecise studies can contribute more noise than useful information.

I also removed the small number of studies with negative effect-size estimates because there were too few to estimate directional selection bias separately. Finally, I excluded samples from non-Western countries. Previous research has shown that psychological effects often vary across countries and cultures, and the present data also suggested substantial country differences. Combining isolated studies from very different populations into a single average would therefore increase unexplained heterogeneity without producing an estimate that clearly applies to either population.

The following analyses therefore focus on a smaller and more homogeneous set of studies and ask a more specific question: How effective is psychotherapy for depression under reasonably comparable conditions in Western countries?

Refined Sample

Random Effects Meta-Analysis

The average effect-size estimate decreased from g = .72 to g = .50. More importantly, heterogeneity decreased dramatically, from tau = .67 to tau = .22. The corresponding 95% prediction interval ranges from approximately g = .06 to g = .94. Thus, although treatment effects still vary substantially, nearly the entire predicted distribution of true effects is now positive.

I next examined study country, waitlist versus treatment-as-usual (TAU) control groups, and type of treatment as moderators. Individual differences associated with country and treatment type were generally modest. The clearest moderator was the control condition: studies using a waitlist produced effects approximately g = .23 larger (SE = .06) than studies using TAU controls.

More important than any individual moderator coefficient, however, was their combined ability to explain heterogeneity. After accounting for country, control condition, and treatment type, residual heterogeneity decreased further to tau = .13. Centered around an overall effect of approximately g = .50, this corresponds to a range of true effects of roughly g = .24 to g = .76.

This is a much more informative answer to the question of psychotherapy effectiveness. Instead of an average surrounded by effects ranging from negative to extremely large, the refined analysis suggests that under reasonably comparable conditions psychotherapy produces effects ranging from about one-quarter to three-quarters of a standard deviation, with an average of about half a standard deviation.

PET

The PET regression again showed a strong relationship between effect-size estimates and their standard errors, b = 2.23, SE = .52. In PET, this relationship is typically interpreted as evidence that effect-size estimates are increasingly inflated as sampling error increases.

However, when country, control condition, and treatment type were added to the regression, the coefficient for sampling error was reduced by about half, from b = 2.23 to b = 1.10, and was no longer statistically significant (SE = .57).

This result illustrates a fundamental limitation of regression-based tests of publication bias. A correlation between effect-size estimates and standard errors does not reveal why the correlation exists. PET attributes this pattern to publication bias, but the same pattern can arise when small and large studies differ systematically in ways that genuinely affect treatment outcomes.

In these data, much of what initially looked like publication bias could be explained by study characteristics. In other words, the apparent bias signal was partly heterogeneity in disguise.

zcurve3

zcurve3 does not yet allow moderators to be specified directly in the model. However, there are two ways to combine zcurve3 with moderator analyses. One approach is to first fit a random-effects meta-regression, remove the variation explained by the moderators, and then analyze the resulting moderator-adjusted effect sizes with zcurve3. To preserve the overall treatment effect, I added the overall mean effect back to the residuals. Thus, the adjusted effects retain the average psychotherapy effect while removing systematic variation associated with country, control condition, and treatment type.

A second approach is to use the bias-corrected study-level effect-size estimates produced by zcurve3 and analyze these estimates in a conventional meta-regression. Here, I used the first approach.

The z-curve plot of the moderator-adjusted effects shows no evidence of excess significance. The observed discovery rate was 79%, almost identical to the expected discovery rate of 78%. If publication selection had produced a large excess of significant findings, the observed discovery rate should have been substantially higher than the rate predicted from the evidential strength of the studies. Instead, the two estimates differ by only one percentage point.

Visual inspection leads to the same conclusion. The fitted z-curve closely reproduces the observed distribution, including the non-significant results to the left of the significance threshold. Thus, after removing systematic variation associated with the moderators, the strong publication-bias signal suggested by PET is no longer apparent.

The overall mean effect is estimated to be g = .53, 95%CI [.09, .59], but the confidence interval remains relatively wide because zcurve3 continues to allow for the possibility of publication selection. In this model, some non-significant results are inferred rather than directly observed, which creates additional uncertainty about the overall mean. If the model were respecified under the assumption that there is no publication bias and all observed results were analyzed directly, the confidence interval would become much narrower—but the analysis would then essentially reduce to another conventional random-effects meta-analysis.

More informative is the convergence of the point estimates across models with very different assumptions. The conventional random-effects model and the selection model both estimate an average effect of approximately half a standard deviation. The mean among the significant results is similarly estimated at g = .52, and the median is g = .54. Heterogeneity is also small, tau = .09, 95% CI [.06, .21] compared to the heterogeneity for all studies (.98, 95%CI [.21, 1.66]. Thus, the smaller dataset does produce a more informative average estimate.

The agreement between the RMA results and these results with a selection model provides an important robustness check. A model that assumes the observed literature is unbiased and a model that explicitly allows for missing non-significant results arrive at essentially the same estimate of psychotherapy effectiveness.

Step-Function Selection Model

To further examine the robustness of the results, I analyzed the data with a step-function selection model (Hedges & Vevea, 1996; Vevea & Hedges, 1995). Unlike PET, this model does not infer publication bias from a correlation between effect sizes and standard errors. Instead, it allows the probability that a result is observed to differ across ranges of p-values.

I specified selection steps at p = .025, corresponding to statistical significance at p = .05 for a two-sided test, and at p = .50, which separates positive from negative effect-size estimates. Because negative effect-size estimates had already been excluded from this analysis, the estimated weight for this final interval is necessarily close to zero and is not substantively informative.

The selection model estimated an average effect of g = .45, SE = .03, with little remaining heterogeneity, tau = .10, 95% CI [.00, .17]. The model did find some evidence of preferential selection of statistically significant results. The relative weight for non-significant positive results was .54, 95% CI [.27, .99]. In other words, the model estimates that non-significant positive results may have been only about half as likely to be observed as statistically significant results, although there is considerable uncertainty about the magnitude of this selection.

Most importantly, allowing for this degree of publication selection had little effect on the estimated treatment effect. Because most results in this refined sample were statistically significant, the estimated mean was largely determined by studies in the region receiving full selection weight.

Thus, a second selection model, based on very different assumptions from zcurve3, reaches essentially the same conclusion: some publication selection may be present, but correcting for it does not materially change the estimated effect of psychotherapy or the amount of remaining heterogeneity.

Robust Bayesian Model Averaging (RoBMA)

A final robustness analysis used Robust Bayesian Model Averaging (RoBMA). RoBMA is useful here because it does not require us to decide in advance whether publication bias should be modeled with PET–PEESE, a step-function selection model, or not at all. Instead, it simultaneously fits a large set of models—including conventional random-effects models, PET–PEESE models, and several step-function selection models—and gives greater weight to models that are better supported by the data.

In the present data, the PET-type models received little support, whereas the selection models received considerably more weight. Thus, when RoBMA was allowed to choose among competing explanations of publication bias, the data favored the step-function approach rather than the PET interpretation.

The posterior probability that the average treatment effect is different from zero was essentially 1.00. There was also strong evidence for remaining heterogeneity, with a posterior inclusion probability of .94, and some evidence for publication bias, with a posterior inclusion probability of .72. These probabilities tell us whether these components are likely to be present, but not how large they are.

RoBMA reports both unconditional estimates, which average across models that include and exclude a particular component, and conditional estimates based only on models in which that component is present. In this analysis, the estimates were essentially identical because the evidence for a treatment effect was overwhelming. The model-averaged effect size was g = .46, SE = .04, and residual heterogeneity was small, tau = .10, 95% interval [.05, .17].

Most importantly, RoBMA independently reproduced the result obtained with the step-function model: some publication selection may be present, but accounting for it leaves the estimated psychotherapy effect at roughly half a standard deviation, with little remaining heterogeneity.

Discussion

What Have We Learned About Meta-Analysis?

This reanalysis illustrates two problems that can make meta-analytic estimates misleading even when sampling error is very small.

First, evidence of a relationship between effect sizes and standard errors should not automatically be interpreted as publication bias. Regression-based methods such as PET can mistake systematic differences between studies for selective publication. In the present data, PET initially suggested that the psychotherapy effect was close to zero. However, the relationship between effect sizes and standard errors was substantially reduced after country, control condition, and treatment type were included as moderators. What initially looked like publication bias was therefore, at least in part, heterogeneity in disguise.

Selection models that do not rely on this correlation produced a very different conclusion. zcurve3 found no clear evidence of excess significance, while the step-function selection model and RoBMA suggested that some publication selection may nevertheless be present. Importantly, correcting for this possible selection had little effect on the estimated treatment effect. Across these different models, estimates converged at approximately half a standard deviation.

The standard recommendation is therefore to use multiple methods and to include selection models. More importantly, discrepancies need to be examined in terms of the different assumptions that models make. Here this analysis showed that one model with different assumptions led to different results because the assumption was false.

Second, meta-analysis should not simply maximize the number of studies and then report heterogeneity as an unfortunate side effect. A precise average of studies that estimate systematically different effects may have little practical meaning. The goal should be to identify moderators that explain these differences and, where necessary, define more homogeneous groups of studies for which an average effect has a meaningful interpretation.

This sometimes means that less is more (Cohen, 1990). A small number of reasonably comparable studies can provide a more useful estimate than a much larger collection of studies that differ substantially in populations, treatments, control conditions, and settings. Unique studies remain scientifically valuable, but a study representing a population or treatment found nowhere else in the meta-analysis cannot tell us whether its unusual effect generalizes beyond that individual study.

The goal is therefore not to eliminate heterogeneity for its own sake. It is to explain heterogeneity well enough that the resulting average describes a meaningful population of studies.

What Have We Learned About Psychotherapy for Depression?

The substantive conclusion is considerably clearer than the original range of meta-analytic estimates from approximately g = .18 to g = .72 suggested.

For reasonably comparable studies conducted in Western countries, psychotherapy produces an average improvement in depression of approximately half a standard deviation compared with treatment as usual. After accounting for country, control condition, and treatment type, remaining heterogeneity was relatively small. The distribution of true study-level effects was approximately g = .2 to g = .8, suggesting that psychotherapy generally produces effects ranging from small to large rather than effects ranging from harmful to extremely large.

The clearest moderator was the control condition. Effects in studies using a waitlist were approximately .2 standard deviations larger than effects in studies comparing psychotherapy with treatment as usual. Country and treatment type also explained some variation, although their individual differences were generally smaller.

Thus, the best answer to the question “How effective is psychotherapy for depression?” is not a single universal number. For the types of studies examined here, a reasonable estimate is about g = .5 compared with treatment as usual, with somewhat larger effects against waitlist controls and modest variation across countries and treatment types.

These findings also point to the limits of further small psychotherapy trials. Small studies were sufficient to establish that psychotherapy works. They are much less useful for determining whether one therapy works slightly better than another or which patients benefit most. Detecting these smaller differences requires much larger samples in which other study characteristics are held reasonably constant.

The next step therefore is not simply to accumulate more small studies. It is to conduct large, coordinated, multi-site studies that can estimate treatment effects precisely enough to determine which treatments work best, under which conditions, and for which patients.

Conclusion

Psychotherapy works, p<.05, is not a scientific conclusion, even when it is based on a meta-analysis of hundreds of studies. Here I showed that the existing evidence allows for a more informative answer, at least for the conditions represented by studies conducted in Western countries. Compared with treatment as usual, psychotherapy reduces depression symptoms by about half a standard deviation on average. This is a clinically meaningful effect, but it is still an average. Across reasonably comparable studies, typical treatment effects appear to range from roughly one-quarter to three-quarters of a standard deviation. Variation across individual patients is likely to be considerably larger.

Thus, how much psychotherapy will help a particular patient remains uncertain. What the evidence does show is that, under the conditions examined here, true study-level effects that favor the control condition appear to be uncommon. Given the substantial average benefit of psychotherapy for patients with depression, psychotherapy should be offered as a core treatment option.

Primed for Equivocation

“Priming exercise” and psychological priming share a label, not necessarily a mechanism. Deliberate preparation for a known future competition provides no obvious need for unconscious goal activation, and the article presents no evidence that such activation actually occurs. The shared terminology leaves the proposed explanation primed for equivocation.

Holmberg and Kelly (2026) provide a useful critique of physiological explanations for “priming exercise,” but their attempt to connect this literature to psychological priming introduces a much less plausible mechanism. The connection appears to arise largely because the two literatures happen to use the same word. That shared terminology leaves the argument primed for equivocation.

In sports physiology, “priming exercise” refers to a deliberately performed bout of exercise intended to improve performance later that day. In cognitive and social psychology, priming refers to prior exposure to a stimulus that subsequently alters processing or behavior, sometimes without awareness or conscious intention. These are fundamentally different uses of the term. The fact that both involve something occurring before something else does not imply that they share a psychological mechanism.

This distinction becomes especially important when the authors invoke nonconscious goal priming. They suggest that exercise might activate performance goals and related behavioral representations that persist until later testing. They cite classic social-psychological priming research, including Bargh and colleagues, to support the possibility that goals can be activated without awareness and subsequently influence behavior.

But the proposed mechanism is poorly matched to the phenomenon being explained. An athlete does not ordinarily encounter a “priming exercise” incidentally. The athlete performs it because a competition or performance test is coming later. The later performance goal is therefore likely to have been activated before the exercise begins:

competition later → intention to prepare → priming exercise.

The goal is not plausibly dormant until the exercise somehow activates it unconsciously. Indeed, the goal is probably one of the reasons the athlete performs the exercise in the first place. During the exercise the athlete may also consciously think about the competition, technique, pacing, readiness, or expected benefits. Under those circumstances, invoking nonconscious goal activation is not merely unnecessary; the proposed causal sequence is almost backwards.

The authors’ own examples illustrate the problem. They discuss athletes rehearsing particular pacing strategies, using metronomes, receiving verbal cues, believing that squatting will improve subsequent jumping, and developing confidence or expectations about later performance. These are readily understood as deliberate preparation, task practice, expectancy, motivation, or attentional effects. None requires a nonconscious priming mechanism.

The scientific evidence offered for the nonconscious account is also weak. The article provides no exercise experiment demonstrating that a priming-exercise bout activates a previously inactive goal outside awareness and that this activation subsequently causes improved performance hours later. In fact, the authors repeatedly acknowledge that these possibilities “may” occur, are “hypothesized,” or “have yet to be directly examined.” Thus, the proposed mechanism is not an empirical finding from the exercise literature.

Instead, support is imported from a different literature on behavioral and goal priming. That literature is itself scientifically controversial, and citing classic demonstrations does not establish that the same mechanism operates in an entirely different situation involving intentional athletic preparation. The conceptual inference appears to be:

psychology calls something “priming”

  • exercise science calls something “priming”
    → psychological priming may explain exercise priming.

But identical terminology is not evidence of mechanistic equivalence.

The distinction matters because the paper already identifies much more plausible explanations. Task-specific practice, motor learning, expectancy, researcher effects, motivation, and explicit performance preparation could all produce later performance changes. These mechanisms fit the actual structure of the situation: athletes know that performance is coming and intentionally prepare for it. The nonconscious goal-priming hypothesis adds an unnecessary and poorly supported causal layer.

The problem can therefore be summarized simply: an implausible mechanism is invoked to explain effects that have not been shown to require that mechanism, and the empirical justification comes largely from a separate literature that happens to use the same word.

Methodological Problems in Claims About “Conscious” and “Preconscious” Influencer Effects

Mir, I. A. (2026). Influencer’s physical attractiveness and content aesthetics: Conscious and preconscious determinants of fashion-branded content engagement on Instagram. Journal of Creative Communications, 21(2), 201–219. https://doi.org/10.1177/09732586241288672

“The authors invoke an unproven perception–behavior mechanism to explain causation that was never observed. A direct path in a cross-sectional SEM establishes neither causation nor preconscious processing.”

Mir (2026) examines whether fashion influencers’ physical attractiveness and the aesthetics of their branded content predict followers’ engagement on Instagram. The study uses survey responses from 300 followers of 15 fashion influencers in Pakistan and analyzes the proposed relationships with structural equation modeling and mediation analyses. The main empirical finding is straightforward: followers who rate influencers and their content more positively also report more favorable attitudes and greater engagement. The methodological problem is that the article draws causal and psychological-process conclusions that the design cannot support.

The most serious problem is that all variables were measured in a single cross-sectional self-report survey. Participants simultaneously reported how attractive they considered the influencer, how aesthetically pleasing they considered the content, their attitude toward that content, and how often they viewed, liked, commented on, and shared it. Nothing was manipulated, and there was no temporal ordering of the variables. Nevertheless, the article repeatedly describes attractiveness and aesthetics as factors that “cause,” “trigger,” “stimulate,” or “activate” engagement. Those causal statements do not follow from the design.

For example, the proposed model assumes

attractiveness → attitude → engagement.

But the same covariance pattern is compatible with numerous alternatives. Followers who engage frequently with an influencer may develop more positive attitudes and subsequently rate that influencer as more attractive. A general liking or identification with the influencer could simultaneously increase attractiveness ratings, content-aesthetic ratings, attitudes, and engagement. The structural equation model cannot distinguish among these explanations.

The sampling procedure makes this problem particularly important. Participants had to have followed one of the selected influencers for more than six months. Thus, the sample is already conditioned on sustained interest in the influencer. People who disliked the influencer, found the content unattractive, or disengaged from it are systematically less likely to appear in the sample. The resulting correlations describe differences among an already selected group of followers; they provide weak evidence for the article’s practical recommendation that firms should hire physically attractive influencers because attractiveness causes engagement.

The article’s central methodological error is even more fundamental. It claims to distinguish a “conscious” route from a “preconscious” route. The mediated path

attractiveness/aesthetics → attitude → engagement

is interpreted as conscious influence, whereas a remaining direct path from attractiveness or aesthetics to engagement is interpreted as evidence of a preconscious perception–behavior process.

A direct regression coefficient is not a measure of unconscious processing.

If attractiveness predicts engagement after statistical adjustment for an explicit attitude measure, this merely shows residual covariance between those variables. That residual association could reflect measurement error in attitude, omitted mediators, stable preferences, common response tendencies, reverse causation, or numerous other processes. Nothing in the study measures awareness, intention, automaticity, processing speed, or participants’ ability to report the causes of their behavior. Consequently, the data provide no evidence that engagement was “preconscious” or unintentional.

For the same reason, the indirect path through an explicit attitude measure does not establish a conscious causal mechanism. Participants consciously completed the attitude questionnaire, but that does not mean the psychological process producing their engagement operated consciously. Statistical mediation and conscious psychological mediation are different concepts.

This problem is especially consequential because the “preconscious” interpretation is one of the article’s principal theoretical contributions. The authors explicitly invoke the perception–behavior literature, including Bargh et al. (1996), to justify the claim that a significant direct path demonstrates automatic behavior. But the current research contains none of the experimental procedures that would be required to test an automatic perception–behavior effect. The SEM therefore cannot adjudicate between conscious and unconscious processes.

The measures themselves also create substantial interpretive problems. On page 209, physical attractiveness is measured with “stylish,” “good looking,” “sexy,” and “elegant.” Content aesthetics is measured with “striking,” “wonderful,” “fascinating,” and “lovely,” while attitude is measured with “pleasant,” “good,” “likeable,” and “my favourite.” These constructs are conceptually and evaluatively intertwined. “Wonderful” and “lovely,” for example, are not narrowly aesthetic judgments, while “stylish” and “elegant” are not purely measures of physical attractiveness. Much of the model may therefore reflect a broad positive-evaluation factor rather than distinct psychological constructs connected by causal pathways.

The observed correlations are consistent with this concern. Physical attractiveness correlates .64 with content aesthetics and .64 with attitude; content aesthetics correlates .66 with attitude and .68 with reported content consumption. Demonstrating discriminant validity with the Fornell–Larcker criterion does not eliminate the possibility that halo effects or general liking strongly influence all of these ratings.

The attempt to dismiss common-method bias is also inadequate. All predictors, mediator variables, and outcomes were obtained from the same respondent at the same time using similar rating formats. The authors test whether several sets of items can be represented by a single latent factor and conclude that poor single-factor fit shows that common-method bias is not important. That conclusion does not follow. Common-method variance does not require every item to load on one factor. Several distinguishable constructs can coexist while correlations among them are inflated by shared method, evaluative consistency, acquiescence, or halo effects.

The proposed “snowball effect” suffers from the same causal problem. The authors find that self-reported consumption behaviors—viewing, reading comments, and liking—predict contribution behaviors such as commenting and sharing, and conclude that consumption gradually causes users to progress toward more active participation. Yet consumption and contribution were measured simultaneously. Someone who frequently comments and shares influencer content almost necessarily also views and consumes that content. A positive cross-sectional association therefore does not demonstrate a temporal progression from passive to active engagement. Testing a snowball process would require longitudinal evidence showing that earlier consumption predicts subsequent increases in contribution.

There may also be an unmodeled dependence problem. The 300 respondents followed one of only 15 macro-influencers. Followers of the same influencer are not necessarily independent observations. Influencers may differ systematically in appearance, production quality, follower demographics, posting frequency, and baseline engagement. Those influencer-level characteristics could generate correlations among respondent ratings. The reported SEM appears to treat all 300 followers as independent rather than accounting for clustering by influencer.

Another reporting issue concerns the engagement scale. The response categories are described as 1 = “very often,” 2 = “often,” 3 = “sometimes,” and 4 = “never.” Thus, larger numerical values indicate less engagement. Yet positive path coefficients are consistently interpreted as greater attractiveness and aesthetics producing greater engagement. The article does not clearly state in the reported method that these scores were reverse-coded. If they were reversed before analysis, that transformation should have been explicitly documented. If they were not, the substantive interpretation of the coefficients would be reversed.

Finally, the statistical success of the model should not be confused with strong evidence for the hypotheses. All five proposed hypotheses are supported, including the weakest direct attractiveness effect, β = .12, t = 2.19. The study was not preregistered, and the sample-size justification consists largely of the statement that N = 300 is sufficient for purposive sampling and structural equation modeling rather than an a priori power analysis tied to the focal effects. Good model-fit indices demonstrate that a specified covariance model can reproduce the observed covariance matrix; they do not establish that the arrows in Figure 2 represent the true causal processes.

The study therefore supports a much narrower conclusion than the article claims. Among existing long-term followers of fashion influencers, positive ratings of influencer attractiveness and content aesthetics are associated with positive attitudes and greater self-reported engagement. That descriptive association is plausible and potentially useful.

The study does not establish that physical attractiveness or content aesthetics cause engagement, that attitude mediates these effects causally, that any residual direct relationship reflects a preconscious perception–behavior mechanism, or that passive engagement develops over time into active contribution.

A suitable experimental design would manipulate influencer attractiveness and content aesthetics independently, randomly assign participants to conditions, measure actual engagement behavior, and include measures capable of testing awareness or automaticity. A longitudinal design would be required to test the proposed consumption-to-contribution “snowball” process. Without such evidence, the article’s strongest psychological claims are interpretations imposed on cross-sectional correlations rather than findings produced by the research design.

Auditing a Poor Audit of Elderly Priming

Costa, T. (2026). The Bayesian audit: Evaluating the proportionality of scientific claims to evidence—a case study on social priming and walking speed. Frontiers in Psychology, 17, 1799078. DOI: 10.3389/fpsyg.2026.1799078


This article applies Bayes’ theorem to one conveniently selected t value and calls the result a ‘Bayesian audit.’ Most readers may simply ignore it because it appeared in Frontiers in Psychology. Those who want a more substantive reason can point to this review.

Costa (2026) introduces a “Bayesian audit,” a six-step framework intended to evaluate whether the strength of scientific claims is proportional to the evidence supporting them. The idea is sensible. Statistical significance does not tell us how strongly we should believe a scientific claim, and surprising claims based on weak evidence deserve particularly careful scrutiny. Costa illustrates the proposed method with one of social psychology’s most famous findings: Bargh, Chen, and Burrows’s (1996) claim that priming college students with words related to old age caused them to walk more slowly afterward.

Unfortunately, the audit itself is problematic. It misrepresents important features of the original study, considers only a fraction of the available evidence, and reduces a question about the magnitude and robustness of an effect to a comparison between a null and an inadequately specified alternative hypothesis.

The first problem is surprisingly basic. Costa describes the original finding as based on a study with approximately t(28) = 2.0 and p ≈ .05. But Bargh et al. actually reported two elderly-priming experiments. In Experiment 2a, the comparison was t(28) = 2.86, p < .01. They then conducted Experiment 2b as a replication and again reported slower walking, t(28) = 2.16, p < .05. Costa appears to approximate the weaker second result while failing to mention the stronger first result or even that the original article contained two studies.

That is an odd starting point for an audit. If the purpose is to reconstruct how much evidence supported the claim in 1996, both original studies should be included.

Costa also incorrectly describes participants as being “subliminally exposed to words related to old age.” They were not. Participants consciously read words while completing a scrambled-sentence task. The claimed unconscious component was that participants supposedly did not realize that the elderly-related words subsequently affected their walking. Bargh et al. themselves explicitly distinguished this procedure from subliminal priming; Experiment 3 of their paper used genuinely subliminal presentation of faces.

This distinction matters because Costa uses the apparent implausibility of unconscious effects on motor behavior to motivate skeptical prior probabilities. One should at least characterize the causal claim correctly before assigning a prior to it.

The treatment of replication evidence is even more problematic.

Costa cites Doyen et al. (2012) and Harris et al. (2013) as subsequent replication attempts. Doyen et al. did replicate the elderly-walking paradigm. In a substantially larger study using automated measurement, they found essentially no priming effect. Their second experiment further suggested that experimenter expectations could influence the result.

Harris et al. (2013), however, did not replicate elderly priming at all. They attempted to replicate Bargh et al.’s 2001 high-performance goal-priming experiments, in which achievement words were supposed to improve performance on a cognitive task. Calling Harris et al. a replication of the elderly-walking finding is simply an error.

More importantly, why is a Bayesian audit conducted in 2026 based primarily on one t statistic from 1996?

There is now a substantial literature on behavioral priming. Dai et al. (2023), for example, meta-analyzed 351 studies and 862 effect sizes and concluded that behavioral priming effects could be detected across a large literature. Conversely, Mac Giolla et al. (2024) examined 70 close replication attempts of 49 social-priming findings. Ninety-four percent produced smaller effects than the originals, only 17% were significant in the predicted direction, and among 52 replications conducted without an original author, none was significant in the original direction; the pooled effect for those independent replications was essentially zero.

These sources do not necessarily settle the question. Meta-analyses themselves can be distorted by publication bias and other forms of selection. But that is precisely why an audit should examine them critically. An audit of a 30-year-old scientific claim should evaluate the accumulated evidence, not simply convert one selected original result into a Bayes factor.

There is an even more fundamental problem with the statistical question Costa asks.

Costa assigns prior probabilities of .05, .10, and .20 to the alternative hypothesis and combines these with an estimated Bayes factor of approximately 3. This yields posterior probabilities of .14, .25, and .43, respectively. The arithmetic is straightforward. The interpretation is not.

Why should the prior probability that the effect exists be .05 or .10?

Costa acknowledges that these values are illustrative rather than derived from an elicitation procedure. But these priors largely determine the conclusion that posterior belief remains low. Starting with a 5% probability and multiplying the prior odds by a Bayes factor of 3 inevitably produces a low posterior probability.

More importantly, what exactly is the hypothesis whose prior probability is 5%?

There is a major difference between these propositions:

elderly-related words have exactly zero effect on walking speed;

elderly-related words have some nonzero effect;

elderly-related words have a psychologically meaningful effect;

elderly priming produces effects of the magnitude originally reported;

automatic stereotype activation reliably produces consequential behavioral changes.

These are not the same hypothesis.

The scientifically interesting issue today is probably not whether the population effect is exactly zero. The effect could be d = .05 or d = .10. Such an effect would make the point null hypothesis technically false while providing little support for the dramatic theoretical interpretation of the original experiments.

This is why effect sizes matter. Bargh et al.’s original studies implied very large effects. Subsequent evidence raises the possibility that the true effect, if it exists at all, is much smaller. A useful audit therefore needs to ask how large the effect is and how precisely it has been estimated—not merely whether H0 or H1 receives the larger Bayes factor.

Costa’s procedure also conflates two different kinds of priors. One is the prior model probability: how likely H1 is relative to H0 before seeing the data. The other is the prior distribution over possible effect sizes within H1. A Bayes factor for a composite alternative necessarily depends on the latter. Yet the article emphasizes sensitivity to prior model probabilities while giving much less attention to the effect-size assumptions used to obtain BF₁₀ ≈ 3.

This creates another problem with Costa’s distinction between “evidence” and “belief.” He describes the Bayes factor as quantifying evidence supplied by the data and posterior probability as combining this evidence with prior belief. But a Bayes factor is not simply a property of the observed data. It depends on the statistical models being compared, including the distribution of effect sizes assumed under the alternative hypothesis.

There is also an internal inconsistency in the treatment of the replication evidence. Costa states that the later replication attempts yielded Bayes factors close to 1 and therefore had little evidential impact. That makes sense: a Bayes factor of 1 leaves prior odds unchanged. Yet the subsequent synthesis says that posterior belief “collapses under replication.” It cannot do both. Replications with BF ≈ 1 cannot cause posterior belief to collapse. To demonstrate such a decline, one would need Bayes factors favoring the null or another competing model and then accumulate this evidence formally.

Publication bias is another conspicuous omission. The article is motivated by the replication crisis and explicitly acknowledges that biased data limit the usefulness of evidential measures. Yet the actual Bayesian calculation treats the published Bargh result as though it were an observation selected independently of statistical significance.

That is unrealistic. A BF of 3 obtained from a randomly selected study and a BF of 3 obtained from a literature in which statistically significant and theoretically exciting findings were preferentially published do not have the same evidential implications. If selection contributed to the replication crisis, an audit of the original published evidence needs to take selection seriously.

Costa also describes the original study’s low statistical power as an additional reason for skepticism. Low power certainly matters because significant results from low-powered studies tend to exaggerate effect sizes, particularly in a selected literature. But sample size has already entered the likelihood used to compute the Bayes factor. Low power is therefore not independent evidence against the hypothesis. The additional concern arises from selection, analytic flexibility, measurement error, and effect-size inflation.

The deeper problem is that Costa reduces a scientific question to H0 versus H1 when several competing explanations exist. Doyen et al.’s work raised experimenter expectancy as one possible explanation. Other possibilities include a genuinely small priming effect, effects restricted to particular conditions or individuals, procedural artifacts, or some combination of these mechanisms. A Bayes factor contrasting an exact-zero model with a generic nonzero-effect model cannot determine which causal explanation is correct.

This is especially important because rejecting H0 would not establish Bargh’s theory. Even convincing evidence for a tiny difference in walking speed would not demonstrate that automatic stereotype activation generally controls overt behavior.

The Bayesian audit is therefore based on a reasonable principle but a poor demonstration. Scientific claims should indeed be proportional to evidence. The problem is that assessing proportionality requires accurately identifying the original evidence, considering the accumulated replication literature, evaluating publication bias, distinguishing statistical from substantive hypotheses, and estimating plausible effect sizes and their uncertainty.

Ironically, the elderly-priming case illustrates the weakness of Costa’s audit more effectively than it illustrates its strengths. A proper audit should not ask merely whether one selected t statistic changes the odds that an effect is exactly zero. It should ask what three decades of evidence tell us about the magnitude, robustness, boundary conditions, and causal interpretation of the phenomenon.

On those questions, uncertainty remains. There may be a small elderly-priming effect. The evidence does not establish that the effect is exactly zero. But neither does the accumulated evidence support taking the spectacular effects reported in 1996 at face value. The important scientific task is to estimate what effect remains after accounting for uncertainty and bias. That requires more than Bayes’ theorem applied to one conveniently chosen t value.

Old Evidence for a Fragile Priming Theory

Przybylinski, E. (2026). Whatever you say: Changing transference-based problem behavior with if–then plans. Self and Identity. Advance online publication. https://doi.org/10.1080/15298868.2026.2613846


Przybylinski (2026) reports two experiments examining whether implementation intentions can prevent problematic behaviors triggered by transference. The theoretical logic is straightforward: subtle resemblance to a significant other is assumed to activate that person’s representation automatically, which can then influence memory, goals, and behavior. An if–then plan is proposed to prevent the activated representation from guiding subsequent behavior.

The principal concern is the credibility of the evidence on which this argument rests.

The article treats automatic behavioral priming as a well-established foundation. For example, it cites Bargh et al. (1996), Bargh et al. (2001), Chartrand and Bargh (1999), and related studies as evidence that contextual cues can automatically trigger overt behavior without awareness or intention. Yet behavioral priming is precisely one of the areas most affected by the replication crisis. The article does not discuss this change in evidential status. Thus, evidence that was considered persuasive when these studies were conducted is largely presented in 2026 as though subsequent replication failures had not occurred.

This matters because the two experiments are themselves products of that earlier research era. The author explicitly states that they were conducted as dissertation research in the late 2000s, before preregistration became standard. Study 1 included only 60 participants, or 20 participants in each of the three strategy conditions, despite testing interaction hypotheses. Study 2 included 47 participants. The manuscript provides a power justification, but because the studies were not preregistered, it is unclear whether the reported power analysis reflects a prospectively specified design decision. The reported “post hoc power” provides little additional information.

The results are remarkably successful. The focal interactions involving behavioral readiness, memory, and overt submissive behavior are all statistically significant in the predicted direction. The reported effects are also very large, frequently exceeding d = 1 and reaching d = 1.69 in Study 1. Across the major focal tests, the success rate is effectively 100%, while average observed power based on the reported effects is roughly 80%.

A perfect success rate is not impossible when power is 80%, but it is more successful than expected. Schimmack (2012) emphasized that unusually high success rates relative to estimated power can indicate that published effect sizes and success rates should not be taken at face value. Here, the outcomes within each experiment are dependent, so a simple excess-success or incredibility calculation would not be appropriate. With only two studies, there is no statistical smoking gun. Nevertheless, the combination of small samples, large effects, multiple significant focal outcomes, and absence of preregistration warrants substantial caution.

The historical timing makes this concern more important. These experiments were conducted approximately 15 years before their publication. During that interval, psychology experienced a replication crisis that directly challenged the credibility of the behavioral-priming literature on which the article relies. Yet the 2026 article does not supplement the old experiments with a contemporary, adequately powered, preregistered replication.

This is particularly striking because the central experiment is readily replicable. The author remains at the same institution where the original research was conducted, and the paradigm requires undergraduate participants rather than an unusually difficult population. A preregistered replication with a substantially larger sample could have provided highly informative evidence about whether the large effects observed in the original dissertation studies survive contemporary scrutiny.

The absence of such a replication changes how the evidence should be interpreted. The results are not invalid merely because they were collected before the replication crisis. But neither the reported effect sizes nor the perfect pattern of statistical success should be treated as reliable estimates of the underlying effects without independent replication.

A Speculative Theory of Replication Failures in Social Psychology: The Empty-Self Metaphor

Klein, J. W., & Swann, W. B., Jr. (2026). Social psychology’s empty-self metaphor and the replication crisis. Perspectives on Psychological Science, 21(2), 138–153. https://doi.org/10.1177/17456916251401849.

Klein and Swann (2026) offer an interesting diagnosis of the replication crisis, but the evidence does not support the strength of their theoretical interpretation.

Their central empirical observation is striking. They coded 41 hypotheses from Many Labs 1 and 2 as either consistent with an “empty-self” metaphor or not. None of the nine “empty-self” hypotheses replicated, whereas 26 of 32 “not-empty-self” hypotheses replicated, a difference of 81 percentage points. Given the small number of empty-self studies, however, the estimate is much less precise than the point estimate suggests; an approximate 95% confidence interval for the difference is about 47 to 91 percentage points. The authors appropriately describe the result as preliminary.

The more serious problem is interpretation. The analysis is correlational. Studies were not randomly assigned to use an “empty-self” theory, and the authors did not code plausible confounding variables that could themselves predict replication success. These include the sample size and statistical power of the original study, the strength of the situational manipulation, whether the manipulation was consciously perceived, the causal proximity between manipulation and outcome, and the prior plausibility of the predicted effect. Their own coding examples illustrate the problem. A gray background affecting support for austerity is coded as “empty self,” whereas paying for a workshop affecting attendance is coded as “not empty self.” The latter is still a situational effect. What differs most obviously is that payment is a strong and behaviorally relevant manipulation, whereas background color is a weak and remote one.

Thus, the empirical result may show that studies proposing large effects of weak, incidental situational manipulations replicate poorly. That is interesting, but it is not the same as showing that studies fail because they neglect an enduring self.

The theoretical explanation is also speculative and at times internally strained. Klein and Swann sometimes treat replication failures as evidence that subtle situational manipulations have little or no effect. In the Many Labs studies, this inference can be justified when very large samples produce narrow confidence intervals around zero. But this cannot be generalized to all failed replications. In other literatures, including elderly priming, the available confidence intervals may still be compatible with small effects. Failure to obtain significance is not equivalent to demonstrating an effect of exactly zero.

At the same time, Klein and Swann suggest that subtle situational effects may depend on the person. They argue that people may respond only to cues to which they are “tuned” and that primes may work only when they connect to existing self-representations. That possibility is not new to priming theory. Priming effects have long been assumed to depend on whether participants possess the relevant stereotype or representation, and some priming studies have explicitly predicted interactions—for example, effects of religious primes that depend on participants’ religiosity.

But this moderation account has an important implication. If a prime affects people for whom it is relevant and has little effect on others, the population-average effect should normally be reduced, not eliminated. Unless one assumes theoretically unusual crossover interactions in which the prime produces effects in opposite directions for different people, sufficiently large studies should still detect a small average effect. For elderly priming, for example, it is easy to imagine that some participants might be more responsive to an elderly stereotype than others. It is much harder to explain why the same prime should make another substantial group walk faster.

This creates a useful empirical question that Klein and Swann do not examine. Their 41 hypotheses should be coded not only as “empty self” or “not empty self,” but also according to whether the original hypothesis predicted a main effect or a Person × Situation interaction. If nearly all of the “empty-self” studies in Many Labs tested simple main effects, that matters because the larger priming literature already contains many moderator and interaction hypotheses. The Many Labs sample would then not represent the full theoretical range of priming research. More broadly, the authors should distinguish studies proposing universal effects from studies explicitly predicting conditional effects.

A stronger analysis would therefore code each study independently for situation strength, awareness of the manipulation, personal relevance, causal distance between manipulation and outcome, original sample size and power, prior plausibility, and whether the prediction was a main effect or an interaction. Only then could one determine whether an “empty-self” construct predicts replication after plausible alternative explanations have been taken into account.

The irony is that Klein and Swann criticize social psychology for building theories on weak evidence, but their own “empty-self” explanation is itself a speculative theory supported by weakly diagnostic data. The empirical pattern—0% versus 81% replication—is interesting. The claim that this pattern is caused by neglect of an enduring self is not established.

On a 1-to-5 scale from speculative theory to theory explaining highly credible phenomena, I would rate the article about 2/5. The phenomenon to be explained is credible: some classes of social-psychological findings replicate poorly. The proposed explanation—that they fail because social psychology adopted an “empty-self” metaphor—remains largely speculative.

A Mostly Speculative Social Identity Theory of Digital Identity

“Theories are a dime a dozen” (Ed Diener, personal communication)

Bingley, W. J., Worthy, P., Wiles, J., & Haslam, S. A. (2026). A social identity theory of digital identity. Perspectives on Psychological Science, 21(4), 346–382. https://doi.org/10.1177/17456916261419813.



Bingley and colleagues (2026) propose a new “social digital identity theory” (SDIT) to explain how social identities operate across online, offline, and hybrid environments. The theory is unusually explicit: the authors formulate 20 propositions concerning identity salience, digital platforms, embodiment, well-being, group functioning, polarization, culture, and other outcomes.

On a 1-to-5 continuum from speculative theory to theories that explain highly credible empirical phenomena (Speculative versus Explanatory Theory: A Review of “Narrative Embodiment” – Replicability-Index), I would give SDIT a rating of 2/5 (see .

This is not a judgment that the theory is uninteresting or implausible. A rating of 2 means that much of the distinctive theory remains speculative and that the empirical findings used to motivate it have not been systematically evaluated for credibility.

There are good features. Bingley et al. clearly distinguish established ideas from novel hypotheses. They explicitly acknowledge that many of their propositions—particularly P4 through P11, P18, and P20—have not previously been tested and require empirical evaluation. Other propositions are extensions of the much older social-identity and self-categorization literature. Thus, they do not claim that all 20 propositions are established facts.

The problem is different. When published studies are cited as empirical support, the authors rarely ask how strong that evidence actually is. There is no systematic consideration of statistical power, independent replication, publication bias, or whether an apparently supportive literature consists mainly of selected significant findings.

Embodiment provides a good example

Proposition 5 states that social-identity salience is shaped by embodiment. To motivate this proposition, the authors cite studies suggesting that heat activates anger concepts, social rejection makes rooms feel colder, keeping secrets feels physically burdensome, fist clenching activates related concepts, and facial-muscle activation influences experience. They then extend this reasoning to virtual embodiment and the Proteus effect.

A particularly revealing example concerns elderly priming.

The authors cite a virtual-reality study by Reinhard et al. in which participants embodied an older or younger avatar. Participants who had embodied the older avatar subsequently walked more slowly during the first part of a walking test than participants who embodied the younger avatar, p = .033. However, these differences disappeared during the second half of the walk.

Bingley et al. describe this result as operating similarly to “classic priming studies in social psychology” and cite Bargh, Chen, and Burrows (1996).

That citation is problematic because Bargh et al.’s famous claim that elderly-related primes cause people to walk more slowly became one of the best-known examples of the replication problems in social psychology. Doyen et al. (2012) conducted a larger replication using automated measurement and failed to reproduce the original effect in their first experiment.

None of this replication history is mentioned.

The result is an evidential chain that looks stronger on paper than it really is:

reported embodiment effects → elderly-avatar study → classic elderly priming → support for an embodiment proposition.

There are several citations, but citations are not the same thing as strong evidence.

The elderly-avatar study may be interesting, and the possibility that virtual embodiment affects subsequent behavior certainly deserves further study. But a marginally significant result that is then linked to a famous but poorly replicated priming finding does not provide strong evidence for a broad theoretical proposition about embodiment and social identity.

The same caution applies to meta-analytic evidence. A meta-analysis can summarize a literature very precisely while still giving a misleading estimate if the underlying literature is affected by publication bias and selective reporting. The important question is therefore not simply whether a theory can cite studies—or even a meta-analysis—that support a proposition. The question is whether the phenomenon itself has been demonstrated with credible evidence.

Why 2 rather than 1?

SDIT deserves more than the lowest score because it is not free-form speculation. It builds on established theoretical traditions, makes explicit predictions, distinguishes some novel claims from older ones, and repeatedly identifies propositions that require future empirical tests.

But it does not deserve a middle or high score because much of the distinctive theory is not yet explaining firmly established empirical phenomena. Moreover, the embodiment section demonstrates that the authors sometimes treat published findings as evidence without critically assessing whether those findings survived the replication crisis.

Thus, the 2/5 rating means:

SDIT is a structured and testable theoretical proposal with some grounding in established research, but much of its distinctive content remains speculative, and the empirical literature used to motivate some propositions is treated too uncritically.

The citation of Bargh et al. (1996) is a particularly clear example. In 2026, a theory article should not invoke elderly priming as supportive evidence without informing readers that the original phenomenon itself remains empirically uncertain.

Individual Differences Do Not Salvage The Priming Train Wreck

Aytürk, E., & Saribay, S. A. (2026). Neglect of the individual as a neglected problem: The relevance of combined idiographic-nomothetic approaches for social psychology today. Self and Identity. https://doi.org/10.1080/15298868.2026.2697941

Aytürk and Saribay (2026) make a useful methodological argument in “Neglect of the individual as a neglected problem.” Social psychology typically averages across people even though the same situation may have different psychological meanings for different individuals. They advocate greater use of personalized stimuli and intensive repeated-measures designs to identify person-specific processes before generalizing across people.

They also suggest that this problem might help explain replication failures in social priming. For example, “elderly” might evoke a frail grandparent for one person and a vigorous retiree for another. Averaging across these individuals could obscure different or even opposing effects. This is a reasonable hypothesis. It is also useful that the authors explicitly acknowledge low statistical power, questionable research practices, and publication bias as established contributors to replication failures.

The problem is that they nevertheless cite Bargh et al. (1996) rather uncritically. The famous finding that activating the elderly stereotype makes people walk more slowly is treated as an example of a potentially heterogeneous priming effect, without informing readers that the original effect became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original effect in their first experiment.

There were earlier attempts to explain the inconsistent findings by invoking individual differences. For example, Cesario et al. (2006) reported that elderly priming slowed participants with relatively positive attitudes toward elderly people but sped up those with negative attitudes. Other studies similarly reported moderator effects. But these small-sample studies do not provide strong evidence that a robust main effect was merely hidden by heterogeneity. Cesario et al.’s study, for example, had only about 67 participants for the relevant analysis and did not obtain a conventionally significant overall priming effect. The evidential burden therefore shifted to an interaction estimated from an even less powerful design.

This is important because interaction effects are generally more difficult to estimate precisely than main effects. Finding a significant personality moderator in a small study does not establish that a failed main effect was really caused by individual differences. The moderator itself needs adequate power and independent replication. A more detailed discussion of these purported replications of elderly priming is available in “Elderly Priming: Did It Ever Work?”

Here Aytürk and Saribay’s methodological proposal may actually provide a better way forward. Intensive repeated measurement can obtain much more information from each participant and therefore can study within-person Person × Situation effects with fewer participants than a conventional design may require. This comes at a cost: many observations per person are needed, and ecological or experience-sampling studies typically sacrifice some of the experimental control available in tightly controlled laboratory experiments. Moreover, any between-person moderator still ultimately depends on having enough individuals.

Thus, the proposal itself is worth pursuing. Low power means that the failed priming literature often provides lack of evidence rather than definitive evidence of absence. Small, conditional, person-specific priming effects remain possible.

But they remain hypotheses to be demonstrated. Individual heterogeneity can explain variation in a real effect; it cannot by itself establish that an unreliable effect is real. A 2026 article explicitly concerned with the replication crisis should therefore not cite Bargh et al. (1996) without acknowledging the replication failures that make the existence of the phenomenon itself uncertain.

Is Subliminal Manipulation Really the Threat?

Smolinski, J., Smolinski, R., Kesting, P., & Kröcher, F. (2026). Ethical guidelines for designing, developing, and deploying AI negotiation agents. Group Decision and Negotiation, 35, 56. https://doi.org/10.1007/s10726-026-10011-2

One of the article’s central psychological concerns is that AI negotiation agents could exploit human vulnerabilities through sublimiminal or unconscious influence. This possibility sounds alarming because it suggests that an AI agent might manipulate a negotiator without the person even realizing that an influence attempt is taking place.

But the psychological evidence cited for this possibility is surprisingly weak.

The authors cite Bargh et al. (1996), one of the best-known studies of unconscious behavioral priming. Its famous elderly-priming experiments reported that participants exposed to words associated with old age subsequently walked more slowly. Yet the article does not mention that this finding later became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original priming effect in their first experiment. The subsequent history of the finding raises further questions about whether elderly priming was ever a reliable phenomenon; I review that evidence in more detail in “Elderly Priming: Did It Ever Work?

The problem is broader than the Bargh study. Subliminal persuasion has never been shown to be a particularly powerful method for changing consequential consumer choices. An older meta-analysis of subliminal advertising estimated an effect of only r = .059 on consumer choice. Later positive studies suggest that effects depend on restrictive boundary conditions, but even these results are not robust. Thus, it remains scientifically unknown if or when subliminal attempts to manipulate people actually work.

This raises an important question for the ethical analysis. Why would an AI negotiation agent need subliminal manipulation at all?

Ordinary persuasion operates in full awareness and can be far more direct. A seller can explicitly frame an offer as prestigious, scarce, safe, innovative, or socially desirable. A brand can persuade consumers that “Apple is cool,” and a consumer may knowingly pay considerably more for an Apple product because that identity has value to them. Nothing has to be flashed below the threshold of awareness. The person can see the message, understand the message, and still be influenced by it.

That form of influence is potentially much more relevant to AI negotiation agents. An AI can learn what arguments appeal to a particular person, emphasize some information rather than other information, frame alternatives strategically, appeal to identity or social norms, create urgency, flatter, build trust, or exploit known preferences. None of these processes requires a controversial theory of unconscious behavioral priming.

This distinction matters because the paper arguably directs attention toward the more sensational but less empirically credible risk. The image of an AI system exploiting subliminal psychological mechanisms is disturbing, but current psychology provides little evidence that such techniques constitute a powerful means of controlling behavior. More ordinary forms of conscious persuasion are both more plausible and potentially more consequential.

The ethical concern is therefore legitimate, but the psychological rationale should be reformulated. Rather than claiming that AI will be able to exploit “unconscious priming far more precisely” than humans, a more defensible concern is that AI may become exceptionally effective at personalized persuasion using psychological processes that operate with the target’s awareness.

The problem with citing Bargh is that it directs attention to irrational fears about subliminal manipulation, when AI may use much more powerful strategies to manipulate people with information that is in plain sight.

Speculative versus Explanatory Theory: A Review of “Narrative Embodiment”

Shimada, S., Tanaka, S., Morioka, S., Rode, G., Roy, J.-M., & Rossetti, Y. (2026). Narrative embodiment: A conceptual framework linking the narrative self and the embodied self. New Ideas in Psychology, 83, 101285. https://doi.org/10.1016/j.newideapsych.2026.101285

Shimada and colleagues (2026) propose “narrative embodiment” as a conceptual framework linking two aspects of the self: the narrative self and the embodied self. Their central proposal is that “character,” borrowed from Paul Ricoeur, functions as an intermediary between narratives about who we are and relatively stable behavioral dispositions. Narratives may therefore alter behavior by changing character, while embodied experiences may feed back into character and ultimately change the narrative self. The authors apply this framework to phenomena ranging from identification with fictional characters and stereotype priming to virtual-reality avatars and rehabilitation.

The article is interesting partly because it raises a broader question about how theoretical articles should be evaluated. Not all theories have the same epistemic status. In particular, it is useful to distinguish speculative theories from explanatory theories.

Speculative and explanatory theories

At one end of a continuum are speculative theories. These theories propose mechanisms or constructs that could potentially explain observations, but the empirical phenomena they are intended to explain have not themselves been firmly established. Freud’s psychoanalytic theories provide a familiar historical example. Concepts such as repression, the id, ego, and superego offered potentially interesting explanations of human behavior, but the empirical foundations for many of these explanations were weak and the theories were sufficiently flexible to accommodate many different observations.

At the other end are theories constructed to explain credible empirical findings. Perception research provides many examples. Psychophysical phenomena can often be demonstrated repeatedly under highly controlled conditions. Researchers may disagree about why a particular perceptual phenomenon occurs, but there is little disagreement that the phenomenon itself occurs. In this situation, theories compete to explain reliable data. The theoretical problem is not “Does this phenomenon really exist?” but “What mechanism explains it?”

This distinction can be represented on a five-point continuum:

  1. Highly speculative — the theory attempts to explain reported findings whose credibility or replicability has not been established.
  2. Mostly speculative — some credible empirical evidence exists, but important parts of the empirical foundation remain uncertain.
  3. Mixed — the theory integrates both well-established findings and substantially less certain findings.
  4. Mostly explanatory — the major empirical phenomena are supported by strong and reasonably replicable evidence.
  5. Explanatory theory of credible findings — the phenomena to be explained are firmly established; the principal uncertainty concerns their explanation rather than their existence.

An important implication of this distinction is that the number of empirical citations in a theoretical article does not necessarily make the theory empirically grounded. A theory can cite dozens of experiments and nevertheless remain highly speculative if those experiments are unreliable. What matters is the credibility of the phenomena that the theory attempts to explain.

Narrative embodiment

On this continuum, I would rate the narrative-embodiment framework 1 out of 5.

This does not mean that the theory is necessarily false. In fact, several aspects of it are intuitively plausible. It seems entirely possible that people’s narratives about themselves influence their behavior and that behavioral experiences subsequently influence how people think about themselves. The authors also make a constructive attempt to translate their framework into empirically measurable variables such as identification, self-efficacy, body image, and behavioral expression.

The problem is that the article does not critically evaluate the credibility of the empirical phenomena used to motivate the theory. Instead, published findings are typically treated as facts that require explanation.

The clearest example appears in the section on stereotype effects. The authors cite the famous study by Bargh, Chen, and Burrows (1996), in which participants exposed to words associated with old age reportedly walked more slowly after leaving the laboratory. Shimada et al. describe the result as follows:

“This demonstrates that stereotypes can shape behavior even without conscious awareness.”

The word “demonstrates” is important. The Bargh finding is not presented as a historically influential but controversial result. It is presented as empirical evidence for the proposed theoretical framework.

Yet the credibility of this finding has been questioned for well over a decade. Doyen et al. (2012) conducted a larger replication using automated measurement of walking speed. Their first experiment found essentially no difference between the elderly-prime and control conditions. The study included 120 participants, compared with 60 across Bargh et al.’s two original elderly-priming experiments, and used infrared sensors rather than manual stopwatch measurements. (PLOS)

Doyen et al.’s second experiment also showed that the outcome depended on experimenters’ expectations. The predicted slowing effect appeared when experimenters had been induced to expect primed participants to walk more slowly. Doyen et al. therefore concluded that priming alone was insufficient to reproduce the original walking-speed effect under their conditions. (PLOS)

Whatever one’s final judgment about behavioral priming, these findings clearly make the empirical status of the Bargh effect relevant to any theoretical review published in 2026. Yet Shimada et al. cite Bargh et al. without mentioning Doyen et al. or the broader controversy surrounding the replicability of behavioral priming.

This is not merely a minor omission in the reference list. It illustrates a fundamental problem in building psychological theories from the published literature. If significant findings are accepted at face value, virtually any collection of published results can provide apparent empirical support for a theoretical framework. After the replication crisis, however, the existence of a published finding cannot be equated with the existence of a credible phenomenon.

A more detailed examination of the history of the elderly-priming effect shows why this distinction matters. The evidence was much less consistent than the textbook version of the finding suggests, and later large preregistered studies provided additional negative evidence. A meta-analysis shows that the existing evidence is so weak and inconsistent that the true effect can range from no effect to a very strong effect without any known moderators (“Elderly Priming: Did It Ever Work?” ReplicabilityIndex)

A theory in search of credible phenomena

The same concern applies more broadly to the article. The authors draw on identification with fictional characters, stereotype effects, the Proteus effect, avatar embodiment, self-efficacy, narrative therapy, and other literatures. Some of these phenomena may ultimately prove more robust than others. Indeed, the authors cite meta-analytic evidence for the Proteus effect and occasionally acknowledge inconsistent findings. But there is no systematic attempt to distinguish highly credible phenomena from literatures characterized by small studies, selective reporting, or replication problems.

Consequently, the empirical literature functions primarily as illustration rather than as a set of established facts that constrain the theory.

This distinction is important because a genuinely explanatory theory should face constraints imposed by reliable observations. Suppose, for example, that repeated high-powered experiments established that identification with a fictional character reliably changed several character-consistent behaviors that had never been directly primed. A theory would then be needed to explain why these effects occur together. Shimada et al.’s concept of “character” might provide one possible explanation.

Indeed, one of the most promising ideas in the article points in precisely this direction. The authors suggest that activating one aspect of a character should affect other behaviors associated with the same character. Experiencing Superman’s ability to fly, for example, might subsequently increase helping behavior or bravery. If such cross-behavior effects were demonstrated reliably, especially under conditions that distinguished character identification from simpler priming or expectancy explanations, the narrative-embodiment framework would become substantially more explanatory and less speculative.

At present, however, the direction of inference is largely reversed. Published findings are assembled into a framework first, while the reliability of those findings is mostly taken for granted.

Conclusion

Narrative embodiment is an imaginative theoretical proposal. It integrates philosophical ideas about narrative identity and embodiment with several areas of psychological research and generates potentially testable hypotheses. For a journal called New Ideas in Psychology, this type of conceptual speculation is entirely appropriate.

But it is important to distinguish an interesting idea from an empirically established explanation.

On a continuum from speculative theory to explanation of credible empirical findings, I would rate the article:

1 / 5 — highly speculative.