Replicability Index: A Blog by Dr. Ulrich Schimmack

Blogging about statistical power, replicability, and the credibility of statistical results in psychology journals since 2014. Home of z-curve, a method to examine the credibility of published statistical results.

Show your support for open, independent, and trustworthy examination of psychological science by getting a free subscription. Register here.

For generalization, psychologists must finally rely, as has been done in all the older sciences, on replication” (Cohen, 1994).

DEFINITION OF REPLICABILITYIn empirical studies with sampling error, replicability refers to the probability of a study with a significant result to produce a significant result again in an exact replication study of the first study using the same sample size and significance criterion (Schimmack, 2017). 

See Reference List at the end for peer-reviewed publications.

Mission Statement

The purpose of the R-Index blog is to increase the replicability of published results in psychological science and to alert consumers of psychological research about problems in published articles.

To evaluate the credibility or “incredibility” of published research, my colleagues and I developed several statistical tools such as the Incredibility Test (Schimmack, 2012); the Test of Insufficient Variance (Schimmack, 2014), and z-curve (Version 1.0; Brunner & Schimmack, 2020; Version 2.0, Bartos & Schimmack, 2021). 

I have used these tools to demonstrate that several claims in psychological articles are incredible (a.k.a., untrustworthy), starting with Bem’s (2011) outlandish claims of time-reversed causal pre-cognition (Schimmack, 2012). This article triggered a crisis of confidence in the credibility of psychology as a science. 

Over the past decade it has become clear that many other seemingly robust findings are also highly questionable. For example, I showed that many claims in Nobel Laureate Daniel Kahneman’s book “Thinking: Fast and Slow” are based on shaky foundations (Schimmack, 2020).  An entire book on unconscious priming effects, by John Bargh, also ignores replication failures and lacks credible evidence (Schimmack, 2017).  The hypothesis that willpower is fueled by blood glucose and easily depleted is also not supported by empirical evidence (Schimmack, 2016). In general, many claims in social psychology are questionable and require new evidence to be considered scientific (Schimmack, 2020).  

Each year I post new information about the replicability of research in 120 Psychology Journals (Schimmack, 2021).  I also started providing information about the replicability of individual researchers and provide guidelines how to evaluate their published findings (Schimmack, 2021). 

Replication is essential for an empirical science, but it is not sufficient. Psychology also has a validation crisis (Schimmack, 2021).  That is, measures are often used before it has been demonstrate how well they measure something. For example, psychologists have claimed that they can measure individuals’ unconscious evaluations, but there is no evidence that unconscious evaluations even exist (Schimmack, 2021a, 2021b). 

If you are interested in my story how I ended up becoming a meta-critic of psychological science, you can read it here (my journey). 

References

Brunner, J., & Schimmack, U. (2020). Estimating population mean power under conditions of heterogeneity and selection for significance. Meta-Psychology, 4, MP.2018.874, 1-22
https://doi.org/10.15626/MP.2018.874

Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17, 551–566
http://dx.doi.org/10.1037/a0029487

Schimmack, U. (2020). A meta-psychological perspective on the decade of replication failures in social psychology. Canadian Psychology/Psychologie canadienne, 61(4), 364–376. 
https://doi.org/10.1037/cap0000246

Mastodon

How Effective is Psychotherapy for the Treatment of Depression?

Conclusion: A meta-analysis of clinical trials in Western nations that compare psychotherapy to treatment as usual shows an average effect size of half a standard deviation with 95% of effects ranging from a quarter to three-quarters of a standard deviation. Thus, psychotherapy is an important component of treatment for depression.

Introduction

Psychotherapy works, p<.05p < .05.

However, a statistically significant result alone does not really help practitioners and patients assess the benefits of psychotherapy. The important question is not “Is the effect greater than zero?” but “How much does psychotherapy help?”

A single study cannot answer this question precisely because psychotherapy studies tend to have relatively small samples and therefore substantial sampling error. Meta-analysis was developed to address this problem. A simple meta-analysis combines the effect-size estimates from individual studies while taking their sampling error (standard errors) into account to obtain a more precise estimate of the average effect. As the number of studies increases, sampling error in the average estimate can become practically negligible.

The headline estimate of psychotherapy effectiveness is about 70% of a standard deviation on a measure of clinical depression such as the Beck Depression Inventory (g=.72g=.72), with very little sampling error (SE=.03SE=.03; Plessen et al., 2023). This suggests that psychotherapy has a substantial positive effect. However, although the average effect can be estimated very precisely, two other sources of uncertainty undermine the usefulness of this estimate: (a) publication bias and (b) variation in true effect sizes across populations, treatments, control conditions, and other study characteristics.

Publication Bias

One problem is publication bias. Studies that show that psychotherapy is effective may be more likely to be published than studies that fail to show an effect. If so, the published literature will exaggerate effectiveness. Plessen et al. (2023) therefore included PET–PEESE, a statistical method designed to estimate the effect after accounting for a possible relationship between effect size and sampling error.

The method exploits the fact that small studies need larger estimated effects to achieve statistical significance, whereas large studies can achieve significance with smaller estimated effects. If statistically significant findings are preferentially published, this creates a relationship between sampling error and observed effect sizes. PET–PEESE uses this relationship to estimate what the effect would be as sampling error approaches zero. Across the PET–PEESE analyses in Plessen et al.’s multiverse, the estimated effect averaged only g=.18g=.18.

The problem is that publication bias is not the only reason why effect sizes might be related to study size. Small and large studies may differ systematically in other ways. For example, smaller studies might provide more intensive and costly treatments, whereas larger trials might use briefer or online interventions. Control groups, patient populations, and other study characteristics may also differ with study size. If these characteristics genuinely influence treatment effects, PET–PEESE can mistake real differences in treatment effectiveness for publication bias and adjust the effect downward too strongly.

We are therefore left with remarkably different answers to a seemingly simple question. A conventional meta-analysis suggests an effect of g=.72g=.72, whereas PET–PEESE suggests an effect closer to g=.18g=.18. Is psychotherapy highly effective, only modestly effective, or somewhere in between?

Fortunately, other methods can help distinguish publication bias from genuine variation in treatment effects. The first aim of this blog post is to use these methods to obtain a more credible estimate of the effectiveness of psychotherapy.

Heterogeneity

The second problem is heterogeneity. The effectiveness of psychotherapy may vary across populations, types of treatment, control conditions, and other study characteristics. In other words, there may be no single effect size that describes the effectiveness of psychotherapy under all conditions. An average can still be calculated, but it may not provide a useful prediction for any particular treatment setting.

In meta-analysis, variation in the true effect sizes across studies is called heterogeneity. In Plessen et al.’s full three-level meta-analysis, the average effect was g = .72, but the between-study variance was tau² = .364, corresponding to tau = .60. Thus, although the average effect was estimated very precisely, the true effects varied substantially from study to study.

Assuming a normal distribution of true effects, approximately 95% of study-level effects would be expected to fall between g = -.46 and g = 1.90. Thus, the same meta-analysis that estimates the average effect very precisely also allows for true effects ranging from moderately favoring the control condition to extremely large benefits of psychotherapy.

This wide range shows why a precisely estimated average does not necessarily answer the question, “How effective is psychotherapy?” An average of g = .72 tells us that psychotherapy is beneficial on average, but it provides little guidance about the effect we should expect in a particular population, treatment, or comparison condition.

To obtain more informative estimates, we need to understand why treatment effects vary across studies. The second aim of this blog post is therefore to identify one or more groups of studies with reasonably similar true effect sizes and to estimate how effective psychotherapy is under more specific conditions.

This goal may seem counterintuitive because a general rule in statistics is that larger samples are more informative. This is clearly true within an individual study: larger samples reduce random sampling error and produce more precise estimates. In meta-analysis, however, simply adding more studies does not necessarily make the answer more informative. Adding studies reduces sampling error in the average effect, but it can also increase heterogeneity if the added studies examine different populations, treatments, countries, or control conditions.

In this sense, less can be more. A meta-analysis of a smaller but more comparable set of studies may provide a more useful estimate than a much larger meta-analysis that averages over systematically different conditions. For example, an estimate of the effectiveness of cognitive behavioral therapy for adults within a particular healthcare setting and relative to a particular control condition may be more informative than a single average that combines different countries, therapies, populations, and comparison groups.

The goal is therefore not to make the meta-analysis as large as possible, but to define groups of studies for which an average effect has a clear and useful interpretation.

Munder et al.’s Meta-Analysis

The Study

I used Munder et al.’s (2022) meta-analysis because it examined a particularly plausible moderator: the treatment received by patients in the control group. In psychotherapy trials, the control condition plays a role similar to the comparison condition in a drug trial. Some patients are assigned to a waitlist and may receive little or no treatment during the study. In other trials, patients in the control group receive treatment as usual (TAU), and psychotherapy is added only for the treatment group. For example, both groups may receive antidepressant medication, while only the treatment group also receives psychotherapy.

Treatment as usual can also vary substantially in intensity and effectiveness. The more effective the treatment received by the control group, the smaller the additional benefit of psychotherapy is likely to be. This means that an average effect size that combines waitlist controls with different forms of treatment as usual may obscure meaningful differences in effectiveness.

Another plausible moderator is the country in which the study was conducted. Even without a specific hypothesis about cultural differences in the treatment of depression, psychological effects often vary across countries and cultures. Country can therefore serve as a useful proxy for differences in culture, healthcare systems, recruitment practices, and other contextual factors that may influence treatment effects.

A third possible moderator is the type of psychotherapy. Although meta-analyses suggest that many forms of psychotherapy are effective, it remains possible that some treatments produce larger effects than others.

The goal is not only to determine whether these moderators predict effect sizes. Even more important is to determine how much heterogeneity remains after accounting for them. If control condition, country, and treatment type explain a substantial portion of the variation across studies, we can obtain more informative estimates for specific sets of conditions—for example, the expected effect of CBT in U.S. studies compared with a waitlist control group.

Random Effects Meta-Analysis

I first analyzed the data with a standard random-effects meta-analysis using the R package metafor. This model estimates the average effect size and the amount of heterogeneity while assuming that the observed studies are not distorted by publication bias. The analysis produced an average effect of g = .72, SE = .08, together with substantial heterogeneity, tau = .67. These results closely replicate the findings of Plessen et al. (2023): psychotherapy appears highly effective on average, but treatment effects vary greatly across studies.

PET

I next applied PET, one of the regression-based methods used by Plessen et al. (2023) to address publication bias. PET tests whether effect-size estimates are related to their standard errors and extrapolates this relationship to a hypothetical study with no sampling error.

The analysis showed a strong relationship between effect sizes and sampling error, b = 3.29, SE = .57. The estimated intercept at zero sampling error was slightly negative, g = -.07, SE = .15, and not significantly different from zero. Taken literally, PET would therefore suggest that there is no convincing evidence for an average psychotherapy effect after correcting for publication bias.

However, this interpretation depends critically on the assumption that the relationship between effect size and sampling error is caused by publication bias. Moreover, substantial heterogeneity remained even after fitting PET, tau = .57. Thus, the model still allows for large positive effects in some studies while simultaneously estimating an average effect close to zero.

zcurve3

To examine publication bias with fewer assumptions about the relationship between effect size and study size, I also analyzed the data with zcurve3. Z-curve converts each effect size and its standard error into a z-value. A two-sided z-value greater than 1.96 is statistically significant. Z-curve uses the distribution of statistical evidence, particularly the significant z-values, to estimate the underlying distribution of evidential strength and to predict how many significant and non-significant results should be observed.

Figure 1 shows that the fitted distribution predicts somewhat more non-significant results than were actually observed. The expected discovery rate—the proportion of statistically significant findings predicted by the model—is only 39%, whereas the observed discovery rate is higher. This pattern is consistent with some selection for statistical significance. However, unlike PET, z-curve is very uncertain about the magnitude of this bias. The 95% confidence interval for the expected discovery rate extends as high as 83%, so the data are also compatible with little or no excess of significant findings.

Z-curve also reveals substantial variation in the strength of evidence across studies. Studies with non-significant z-values have low estimated power, whereas many of the statistically significant studies have moderate to high power, ranging from approximately 62% to 98%. The average estimated power of the significant studies—the Expected Replication Rate—is 74%. Z-curve also estimates the maximum proportion of statistically significant findings that could be false positives. Although the upper bound of the 95% confidence interval reaches 42%, even this conservative bound implies that the majority of statistically significant findings are unlikely to be false positives.

Zcurve3 can also provide estimates on the effect-size scale. The estimated overall mean effect is g = .33, with a wide 95% confidence interval ranging from g = .12 to g = .90. This interval contains both the small PET estimate and substantially larger conventional random-effects estimates. Rather than forcing the data toward one of these answers, zcurve3 makes the uncertainty about publication bias explicit, rather than assuming that bias is large (PET) or that there is no bias (RMA).

Making Sense of Heterogeneity

The z-curve plot shows heterogeneity in the strength of evidence below the x-axis. These values are estimates of local statistical power for studies with z-values in the corresponding ranges. Studies with small z-values have low estimated power and therefore provide little information about the magnitude of the underlying true effect. In the present data, local power is below 50% throughout the non-significant range. Around z = 2, however, estimated local power rises above 50%, reaching 62% for results just above the conventional significance thresho ld and increasing further for larger z-values.

This provides a principled way to distinguish relatively informative from highly uncertain effect-size estimates. Importantly, the criterion is not statistical significance itself. The criterion is estimated local power. In these data, the point at which local power exceeds 50% happens to coincide approximately with the conventional significance threshold. Thus, focusing on the statistically significant results in this particular dataset amounts to focusing on the subset for which zcurve3 estimates that there is more signal than noise.

zcurve3 also estimates the maximum false-positive rate. The upper bound of the 95% confidence interval is 39%. Thus, even under a conservative interpretation, the majority of results in this more informative subset are estimated to reflect a genuine positive treatment effect.

The main advantage of zcurve3 is that it is designed to examine heterogeneity. There are two types of heterogeneity to consider. Heterogeneity in the strength of evidence and heterogeneity in effect sizes.

The z-curve plot shows heterogeneity in the strength of evidence below the x-axis. These values are local power estimates for the corresponding ranges of z-values. Average power is low for non-significant results. These are mostly studies with small samples and large sampling error, and they provide little information about the magnitude of the underlying true effects. In contrast, studies with z-values greater than about 2 have considerably greater evidential strength. For results just above the conventional significance threshold, estimated local power is already 62% and increases further for larger z-values.

zcurve3 also estimates the maximum false-positive rate. The upper limit of the 95% confidence interval is 39%. Thus, even under this conservative estimate, the majority of the significant results are expected to reflect a genuine positive treatment effect. It is therefore possible to identify a subset of studies that provides substantially stronger evidence about treatment effectiveness.

zcurve3 also provides effect-size estimates for subsets of studies defined by their observed z-values. For the statistically significant results, the estimated mean effect size is g = 1.41, but uncertainty is substantial, 95% CI [.56, 2.21]. Heterogeneity among these effects is also very large, tau = .98, with a wide 95% CI [.21, 1.67]. Thus, focusing on studies with stronger evidence does not solve the heterogeneity problem. We still need to ask why some studies produce much larger effects than others. This requires examining potential moderators—that is, study characteristics that explain variation in effect sizes—with the goal of identifying one or more groups of studies for which an average effect size has a meaningful interpretation.

It therefore makes sense to examine studies with strong evidence more closely. A complication is that unusually large observed effects may partly reflect sampling error or selection. zcurve3 addresses this problem by using empirical-Bayes shrinkage to produce adjusted effect-size estimates that pull unusually large observed effects toward more plausible underlying values. Figure 2 shows these adjusted estimates.

The first five effects, representing four studies, still appear unusually large even after this adjustment. These studies may be scientifically interesting and deserve careful examination and replication, but they are not necessarily informative about the average effect in the broader group of studies. Instead, if their unusually large effects arise from study characteristics that are not shared by the remaining studies, combining them with the rest simply increases unexplained heterogeneity.

Chiang et al. provides a useful example. It is the only study from Taiwan in this meta-analysis. Consequently, its exceptionally large effect is perfectly confounded with the individual study: with only one Taiwanese study, we cannot determine whether the effect reflects Taiwan, some other feature of the study, or sampling variation. The study is informative about that particular Taiwanese sample, but it cannot establish a general Taiwanese treatment effect and it contributes little to estimating the effect for a population of Western studies. For that purpose, including it mainly adds heterogeneity that cannot be explained or generalized.

To produce a more homogeneous and interpretable set of studies, I applied several additional restrictions. First, I removed five studies with unique characteristics or questionable reporting that made their unusually large effects difficult to interpret or generalize. Second, I removed studies with large sampling error (SE > .30). This is actually a relatively modest restriction compared with Stanley, Jarrell, and Doucouliagos (2010), who proposed estimating meta-analytic effects from only the most precise 10% of published estimates when publication selection is a concern. Their argument is that highly imprecise studies can contribute more noise than useful information.

I also removed the small number of studies with negative effect-size estimates because there were too few to estimate directional selection bias separately. Finally, I excluded samples from non-Western countries. Previous research has shown that psychological effects often vary across countries and cultures, and the present data also suggested substantial country differences. Combining isolated studies from very different populations into a single average would therefore increase unexplained heterogeneity without producing an estimate that clearly applies to either population.

The following analyses therefore focus on a smaller and more homogeneous set of studies and ask a more specific question: How effective is psychotherapy for depression under reasonably comparable conditions in Western countries?

Refined Sample

Random Effects Meta-Analysis

The average effect-size estimate decreased from g = .72 to g = .50. More importantly, heterogeneity decreased dramatically, from tau = .67 to tau = .22. The corresponding 95% prediction interval ranges from approximately g = .06 to g = .94. Thus, although treatment effects still vary substantially, nearly the entire predicted distribution of true effects is now positive.

I next examined study country, waitlist versus treatment-as-usual (TAU) control groups, and type of treatment as moderators. Individual differences associated with country and treatment type were generally modest. The clearest moderator was the control condition: studies using a waitlist produced effects approximately g = .23 larger (SE = .06) than studies using TAU controls.

More important than any individual moderator coefficient, however, was their combined ability to explain heterogeneity. After accounting for country, control condition, and treatment type, residual heterogeneity decreased further to tau = .13. Centered around an overall effect of approximately g = .50, this corresponds to a range of true effects of roughly g = .24 to g = .76.

This is a much more informative answer to the question of psychotherapy effectiveness. Instead of an average surrounded by effects ranging from negative to extremely large, the refined analysis suggests that under reasonably comparable conditions psychotherapy produces effects ranging from about one-quarter to three-quarters of a standard deviation, with an average of about half a standard deviation.

PET

The PET regression again showed a strong relationship between effect-size estimates and their standard errors, b = 2.23, SE = .52. In PET, this relationship is typically interpreted as evidence that effect-size estimates are increasingly inflated as sampling error increases.

However, when country, control condition, and treatment type were added to the regression, the coefficient for sampling error was reduced by about half, from b = 2.23 to b = 1.10, and was no longer statistically significant (SE = .57).

This result illustrates a fundamental limitation of regression-based tests of publication bias. A correlation between effect-size estimates and standard errors does not reveal why the correlation exists. PET attributes this pattern to publication bias, but the same pattern can arise when small and large studies differ systematically in ways that genuinely affect treatment outcomes.

In these data, much of what initially looked like publication bias could be explained by study characteristics. In other words, the apparent bias signal was partly heterogeneity in disguise.

zcurve3

zcurve3 does not yet allow moderators to be specified directly in the model. However, there are two ways to combine zcurve3 with moderator analyses. One approach is to first fit a random-effects meta-regression, remove the variation explained by the moderators, and then analyze the resulting moderator-adjusted effect sizes with zcurve3. To preserve the overall treatment effect, I added the overall mean effect back to the residuals. Thus, the adjusted effects retain the average psychotherapy effect while removing systematic variation associated with country, control condition, and treatment type.

A second approach is to use the bias-corrected study-level effect-size estimates produced by zcurve3 and analyze these estimates in a conventional meta-regression. Here, I used the first approach.

The z-curve plot of the moderator-adjusted effects shows no evidence of excess significance. The observed discovery rate was 79%, almost identical to the expected discovery rate of 78%. If publication selection had produced a large excess of significant findings, the observed discovery rate should have been substantially higher than the rate predicted from the evidential strength of the studies. Instead, the two estimates differ by only one percentage point.

Visual inspection leads to the same conclusion. The fitted z-curve closely reproduces the observed distribution, including the non-significant results to the left of the significance threshold. Thus, after removing systematic variation associated with the moderators, the strong publication-bias signal suggested by PET is no longer apparent.

The overall mean effect is estimated to be g = .53, 95%CI [.09, .59], but the confidence interval remains relatively wide because zcurve3 continues to allow for the possibility of publication selection. In this model, some non-significant results are inferred rather than directly observed, which creates additional uncertainty about the overall mean. If the model were respecified under the assumption that there is no publication bias and all observed results were analyzed directly, the confidence interval would become much narrower—but the analysis would then essentially reduce to another conventional random-effects meta-analysis.

More informative is the convergence of the point estimates across models with very different assumptions. The conventional random-effects model and the selection model both estimate an average effect of approximately half a standard deviation. The mean among the significant results is similarly estimated at g = .52, and the median is g = .54. Heterogeneity is also small, tau = .09, 95% CI [.06, .21] compared to the heterogeneity for all studies (.98, 95%CI [.21, 1.66]. Thus, the smaller dataset does produce a more informative average estimate.

The agreement between the RMA results and these results with a selection model provides an important robustness check. A model that assumes the observed literature is unbiased and a model that explicitly allows for missing non-significant results arrive at essentially the same estimate of psychotherapy effectiveness.

Step-Function Selection Model

To further examine the robustness of the results, I analyzed the data with a step-function selection model (Hedges & Vevea, 1996; Vevea & Hedges, 1995). Unlike PET, this model does not infer publication bias from a correlation between effect sizes and standard errors. Instead, it allows the probability that a result is observed to differ across ranges of p-values.

I specified selection steps at p = .025, corresponding to statistical significance at p = .05 for a two-sided test, and at p = .50, which separates positive from negative effect-size estimates. Because negative effect-size estimates had already been excluded from this analysis, the estimated weight for this final interval is necessarily close to zero and is not substantively informative.

The selection model estimated an average effect of g = .45, SE = .03, with little remaining heterogeneity, tau = .10, 95% CI [.00, .17]. The model did find some evidence of preferential selection of statistically significant results. The relative weight for non-significant positive results was .54, 95% CI [.27, .99]. In other words, the model estimates that non-significant positive results may have been only about half as likely to be observed as statistically significant results, although there is considerable uncertainty about the magnitude of this selection.

Most importantly, allowing for this degree of publication selection had little effect on the estimated treatment effect. Because most results in this refined sample were statistically significant, the estimated mean was largely determined by studies in the region receiving full selection weight.

Thus, a second selection model, based on very different assumptions from zcurve3, reaches essentially the same conclusion: some publication selection may be present, but correcting for it does not materially change the estimated effect of psychotherapy or the amount of remaining heterogeneity.

Robust Bayesian Model Averaging (RoBMA)

A final robustness analysis used Robust Bayesian Model Averaging (RoBMA). RoBMA is useful here because it does not require us to decide in advance whether publication bias should be modeled with PET–PEESE, a step-function selection model, or not at all. Instead, it simultaneously fits a large set of models—including conventional random-effects models, PET–PEESE models, and several step-function selection models—and gives greater weight to models that are better supported by the data.

In the present data, the PET-type models received little support, whereas the selection models received considerably more weight. Thus, when RoBMA was allowed to choose among competing explanations of publication bias, the data favored the step-function approach rather than the PET interpretation.

The posterior probability that the average treatment effect is different from zero was essentially 1.00. There was also strong evidence for remaining heterogeneity, with a posterior inclusion probability of .94, and some evidence for publication bias, with a posterior inclusion probability of .72. These probabilities tell us whether these components are likely to be present, but not how large they are.

RoBMA reports both unconditional estimates, which average across models that include and exclude a particular component, and conditional estimates based only on models in which that component is present. In this analysis, the estimates were essentially identical because the evidence for a treatment effect was overwhelming. The model-averaged effect size was g = .46, SE = .04, and residual heterogeneity was small, tau = .10, 95% interval [.05, .17].

Most importantly, RoBMA independently reproduced the result obtained with the step-function model: some publication selection may be present, but accounting for it leaves the estimated psychotherapy effect at roughly half a standard deviation, with little remaining heterogeneity.

Discussion

What Have We Learned About Meta-Analysis?

This reanalysis illustrates two problems that can make meta-analytic estimates misleading even when sampling error is very small.

First, evidence of a relationship between effect sizes and standard errors should not automatically be interpreted as publication bias. Regression-based methods such as PET can mistake systematic differences between studies for selective publication. In the present data, PET initially suggested that the psychotherapy effect was close to zero. However, the relationship between effect sizes and standard errors was substantially reduced after country, control condition, and treatment type were included as moderators. What initially looked like publication bias was therefore, at least in part, heterogeneity in disguise.

Selection models that do not rely on this correlation produced a very different conclusion. zcurve3 found no clear evidence of excess significance, while the step-function selection model and RoBMA suggested that some publication selection may nevertheless be present. Importantly, correcting for this possible selection had little effect on the estimated treatment effect. Across these different models, estimates converged at approximately half a standard deviation.

The standard recommendation is therefore to use multiple methods and to include selection models. More importantly, discrepancies need to be examined in terms of the different assumptions that models make. Here this analysis showed that one model with different assumptions led to different results because the assumption was false.

Second, meta-analysis should not simply maximize the number of studies and then report heterogeneity as an unfortunate side effect. A precise average of studies that estimate systematically different effects may have little practical meaning. The goal should be to identify moderators that explain these differences and, where necessary, define more homogeneous groups of studies for which an average effect has a meaningful interpretation.

This sometimes means that less is more (Cohen, 1990). A small number of reasonably comparable studies can provide a more useful estimate than a much larger collection of studies that differ substantially in populations, treatments, control conditions, and settings. Unique studies remain scientifically valuable, but a study representing a population or treatment found nowhere else in the meta-analysis cannot tell us whether its unusual effect generalizes beyond that individual study.

The goal is therefore not to eliminate heterogeneity for its own sake. It is to explain heterogeneity well enough that the resulting average describes a meaningful population of studies.

What Have We Learned About Psychotherapy for Depression?

The substantive conclusion is considerably clearer than the original range of meta-analytic estimates from approximately g = .18 to g = .72 suggested.

For reasonably comparable studies conducted in Western countries, psychotherapy produces an average improvement in depression of approximately half a standard deviation compared with treatment as usual. After accounting for country, control condition, and treatment type, remaining heterogeneity was relatively small. The distribution of true study-level effects was approximately g = .2 to g = .8, suggesting that psychotherapy generally produces effects ranging from small to large rather than effects ranging from harmful to extremely large.

The clearest moderator was the control condition. Effects in studies using a waitlist were approximately .2 standard deviations larger than effects in studies comparing psychotherapy with treatment as usual. Country and treatment type also explained some variation, although their individual differences were generally smaller.

Thus, the best answer to the question “How effective is psychotherapy for depression?” is not a single universal number. For the types of studies examined here, a reasonable estimate is about g = .5 compared with treatment as usual, with somewhat larger effects against waitlist controls and modest variation across countries and treatment types.

These findings also point to the limits of further small psychotherapy trials. Small studies were sufficient to establish that psychotherapy works. They are much less useful for determining whether one therapy works slightly better than another or which patients benefit most. Detecting these smaller differences requires much larger samples in which other study characteristics are held reasonably constant.

The next step therefore is not simply to accumulate more small studies. It is to conduct large, coordinated, multi-site studies that can estimate treatment effects precisely enough to determine which treatments work best, under which conditions, and for which patients.

Conclusion

Psychotherapy works, p<.05, is not a scientific conclusion, even when it is based on a meta-analysis of hundreds of studies. Here I showed that the existing evidence allows for a more informative answer, at least for the conditions represented by studies conducted in Western countries. Compared with treatment as usual, psychotherapy reduces depression symptoms by about half a standard deviation on average. This is a clinically meaningful effect, but it is still an average. Across reasonably comparable studies, typical treatment effects appear to range from roughly one-quarter to three-quarters of a standard deviation. Variation across individual patients is likely to be considerably larger.

Thus, how much psychotherapy will help a particular patient remains uncertain. What the evidence does show is that, under the conditions examined here, true study-level effects that favor the control condition appear to be uncommon. Given the substantial average benefit of psychotherapy for patients with depression, psychotherapy should be offered as a core treatment option.

It is OK to be a WEIRD science

It Is Fine to Be WEIRD

Americans love acronyms, and none has traveled further in social psychology than WEIRD — Western, Educated, Industrialized, Rich, and Democratic — coined to criticize a discipline for building a science of humanity out of American undergraduates. The critique was fair, but it points in the wrong direction. Being WEIRD is not the problem. It is perfectly legitimate for WEIRD researchers, paid by WEIRD institutions, to study WEIRD people in order to help WEIRD societies — to ask whether a therapy relieves depression in a Western clinic even if it would do nothing in another culture, or even if the disorder as we define it barely exists there. Local knowledge is not lesser knowledge.

The problem is not being WEIRD. It is being WEIRD while claiming to be universal — and psychology keeps making that claim because it has never quite decided whether it is a natural science of universal laws or a social science of particular societies. This essay is about that confusion, the two rituals that keep it in place — the apologetic limitations paragraph and the meta-analytic average that pools everyone into a number belonging to no one — and what becomes possible once we accept that it is fine to be WEIRD.

Philosophy with p-values

Psychology was born from philosophy but wanted to be physics. It took philosophy’s questions — how we perceive, learn, remember, decide — and set out to answer them with the methods of the natural sciences: experiment, measurement, and laws that hold for everyone. Call it philosophy with p-values. For the questions it started with, this worked reasonably well. Basic perception really is close to universal. A Weber fraction measured in Leipzig is likely to look much the same in Toronto; basic visual processes do not change dramatically from one culture to another. In the laboratory of basic processes, one human is often interchangeable enough with another that findings can generalize broadly.

That success was a trap. It made universality the price of admission to scientific psychology: if your findings held for everyone, everywhere, you were doing real science; if they did not, you were doing something softer and less scientific. The standard was manageable while psychology studied processes that are, in fact, close to universal. It became a problem when psychologists turned to everything else — love, prejudice, persuasion, the self — and kept the same standard.

There are two ways to apply a universal method to a variable subject. The first is to deny much of the subject matter. Behaviorism largely did that: mind, meaning, and emotion were pushed aside in favor of observable stimulus and response, which could be studied in a more law-like fashion. Whatever would not fit the method was ruled unscientific and shown the door.

The second way survived behaviorism and is the one we still practice. Instead of changing the questions, we changed the way we studied them. Social psychology adopted the model of the experimental laboratory — controlled experiments, manipulations, deception, participants reduced to “subjects” — so that social life could be studied as though it, too, obeyed general laws.

What psychology was reluctant to consider was another possibility: that the study of social behavior might be a different kind of science, with standards of its own and no need to discover universal laws.

Biology shows that this was a choice, not a necessity. It is a natural science with genuinely universal principles, and no one doubts its scientific credentials. Yet while reproduction is universal, how organisms reproduce varies enormously, and it would be absurd to insist on one detailed theory of sexual behavior spanning salmon, praying mantises, and swans. Biology does not apologize for this. It allows general evolutionary principles to coexist with the study of particular species and ecological contexts. The second does not have to dissolve into the first to count as science.

The False Promise of Experimental Social Psychology

The first challenge to universality came from studies within WEIRD cultures showing that people behave differently in the same situation. Walter Mischel’s Personality and Assessment (1968) argued that behavior could not be predicted very well from broad personality traits. Social psychologists pushed this argument further. Ross (1977) called the tendency to explain behavior in terms of personality while underestimating the situation the fundamental attribution error, and Ross and Nisbett (1991) went so far as to write that “one cannot predict with any accuracy how particular people will respond.” In this way, the search for universal laws of behavior could continue. Personality psychologists marshalled evidence that people really do differ from one another, but experimental social psychology did not respond by making those differences central to its theories. Social psychologists went on manipulating situations in the laboratory and largely treating differences between people as error variance.

The second challenge came from outside the Western laboratory. Beginning in the 1970s and gathering force through the 1980s, cross-cultural psychologists showed that many supposedly universal findings varied systematically from one culture to another — that the mind studied in Michigan was not simply the human mind. In principle, the point was conceded: virtually everyone now agrees that culture shapes cognition and behavior. In practice, much less changed, because studies kept being run in the same few places.

The complaint finally crystallized, decades later, in Henrich, Heine, and Norenzayan’s 2010 article on the “weirdest people in the world” — WEIRD, for Western, Educated, Industrialized, Rich, and Democratic. If anything, WEIRD may make psychology sound more diverse than it was. France is WEIRD, Germany is WEIRD, Sweden is WEIRD, but the standard participant was not drawn from a representative cross-section of Western societies. The empirical base was often much narrower — closer to a WASP monoculture of white, middle-class American college students.

The article was cited everywhere and changed surprisingly little, much like the cross-cultural critique before it. Papers still open with theories stripped of cultural context and test them on WEIRD samples; they simply close, now, with a limitations paragraph acknowledging that the sample was WEIRD and urging further research on other populations — by someone else. The acknowledgment goes in the one section of a paper that rarely changes the interpretation of the findings, and the work proceeds much as before.

And the asymmetry hides in plain sight, even in the names. Other countries mark their journals — the British Journal of this, the Iranian Journal of that — while the journals that set the field’s agenda do not: they are simply Psychological Science or the Journal of Personality and Social Psychology, as if nationality did not apply to them. So an American study in an American journal is readily read as a finding about people, while an Iranian study in an Iranian journal is read as a finding about Iranians. The universality we claim to have renounced still runs underneath — the default setting for one population, and the denial of that status to every other.

So the apology is the wrong response to the right observation. To acknowledge that a sample was WEIRD and then generalize anyway leaves the universalist assumption intact and merely adds a note of caution. The alternative is not to apologize but to specify — to name the population you studied and claim nothing beyond it. There is nothing unscientific about a finding that holds in one society and not another. Whether a result generalizes across cultures is a question to be answered, not a box to be ticked or an assumption to be smuggled in.

Everybody knows that people differ from one another; that does not make anyone weird. What is truly weird is a science of human behavior that ignores this diversity and imagines that behavior can be reduced to a single number — one universal effect of a situation on everyone. The remedy is not complicated: stop making global generalizations and ask instead a specific question about a specific population. Sometimes, in other words, it is perfectly OK to ask a WEIRD question.

Demographic Acronym “WEIRD” Overused in Psychology Research | Psychology Today

The Weirdest Statistical Method: Meta-Analysis

It is widely recognized that many psychological studies cannot provide conclusive evidence for an effect, let alone against one. Sample sizes are often too small to show that a result is more than a statistical fluke, or to pin down how large an effect actually is. The proposed solution is meta-analysis: find all the published studies on a topic and combine them to estimate the average effect size. That average is then presented as the true population effect size.

Pooling does solve the problem it was built for. Combine enough studies and the average becomes precise — the sampling error that plagued each small study shrinks toward zero. But a precise average is worth nothing if the quantity it averages over is not one quantity at all. If the true effect varies from study to study — across populations, procedures, and cultures — then pooling delivers an exact estimate of a number that describes no one. And heterogeneous effects are exactly what psychology’s meta-analyses pool.

To see why that is a mistake, and to see it clearly enough that no one can accuse me of being against averages, forget psychology for a moment and consider the average height of 8 billion human beings on this planet. Suppose it is 165 centimeters. There is nothing wrong with the number. It is correct, it is precisely estimated, and it is the honest answer to a well-defined question: what is the mean of the height distribution over all living people? Now try to use it. Design a doorframe? You do not build to the mean; you build to a high percentile, because the mean is silent about the tail and the tail is the entire point. Manufacture clothing? Here the average is not merely useless but actively misleading, because no one is average on every dimension at once, and a garment cut to the mean neck, mean arm, and mean torso fits no actual body.

So the average can be correct, precise, and useless, all at once, with no statistical error anywhere. The uselessness is not a flaw in the estimate. It is a property of the question. And the question carries a hidden assumption — one no one would defend if asked, and that the practice acts on regardless. Put the claim baldly and it is absurd: that there is a single true value and every individual simply equals it, the variation being error. No one believes that. but in a psychological meta-analysis the same assumption slips through, not as a belief anyone holds but as a convention everyone follows: the pooled effect is written up as the effect, cited as the effect, carried into the next study as the effect — as though the number applied to everybody.

So, to summarize, estimating a single average for a heterogeneous set of objects is a weird question that no one would consider meaningful to ask or to answer. But its weirdness is hidden by the universality assumption — that variation between people is mere error variance, and that the truth is a single number applying to everybody. A meta-analysis of mindfulness therapy illustrates that I am not attacking a strawman, but that the problem is real. I picked this meta-analysis because I am interested in the effectiveness of mindfulness therapy for WEIRD people in Canada. I don’t want to generalize to all other meta-analyses, but it is likely that this is not the only meta-analysis that failed to take cultural differences into account.

Does Mindfulness Therapy Work?

Goldberg and colleagues (2018) set out to answer exactly that question with a meta-analysis of mindfulness-based therapy. They collected studies spanning a range of disorders — depression, anxiety, substance use, and more — conducted in different populations and countries and using different mindfulness protocols, from standardized programs like MBSR and MBCT to local adaptations. They then pooled these studies to estimate average effects. For depression compared with passive controls, for example, the estimated effect was d =.6, with a 95% confidence interval from .5 to .7 — a moderate effect, the kind of number that easily becomes “mindfulness works” by the time it reaches a textbook or a clinician.

But look at the question that average answers — “the effect of mindfulness therapy” — and you will recognize the problem from the previous section. It treats “mindfulness therapy,” without further qualification, as though there were a meaningful effect to be estimated across all of these studies. Yet the studies were not one thing. Mindfulness for chronic pain in an American clinic and mindfulness for depression in a Chinese university are no more the same treatment of the same disorder in the same population than a newborn and an adult are the same height. The pool is heterogeneous by construction — across disorders, populations, and therapies at once. Asking for its single average effect is like asking for the average height of everyone on Earth.

I reanalyzed the open data with z-curve3 (Schimmack, 2026), which is designed to model heterogeneous evidence. The first step is to check for publication bias, and there was little evidence of it — unusual in psychological research, but less surprising in a meta-analysis that includes many nonsignificant results. This means that the full set of 214 positive effect-size estimates can be analyzed without a large correction for selective reporting. The studies were, on average, modestly powered: only about 45% reached significance, reflecting a literature that mixes a few well-powered studies with many underpowered ones.

But the decisive quantity is not the average power or even the average effect. It is the spread. z-curve3 estimates not only the mean true effect but also the distribution of true effects across studies. The estimated mean was .47, reassuringly close to Goldberg’s pooled estimate, with a standard deviation of .29. Those numbers imply a 95% prediction interval from about −.10 to 1.10. In other words, the true effect in another study drawn from this literature could plausibly range from a negligible negative effect to an enormous positive one.

That interval is the data refusing the question. If “the effect of mindfulness therapy” were a single useful quantity, the studies would cluster around it and the interval would be narrow. Instead, it spans almost the entire range of plausible effects.

And this is the point where the argument is often lost, so I want to be precise. The average of .47 is not meaningless. It is the correct answer to a narrow and legitimate question: if you drew another study at random from this same mixture of studies, .47 would be your best single guess for its effect. But patients are not looking to enter a lottery whose prize ranges from a small harm to a large benefit. They want to know whether a particular therapy has been shown to work for their problem and in a population like theirs.

There are, in fact, two problems stacked on top of each other. Even within a single population, an average conceals variation between individuals — some patients improve, some do not. I set that problem aside here because the meta-analytic average fails long before we reach it. It is already an average of different population averages, different treatments, and different disorders. It tells us little about how any particular therapy performs for any particular problem in any particular population. The within-population question is hard. This broader question may not even be well posed.

Goldberg and colleagues (2018) set out to answer exactly that question with a meta-analysis of mindfulness-based therapy. They collected studies spanning a range of disorders — depression, anxiety, substance use, and more — conducted on different populations in different countries, using different mindfulness protocols, from standardized programs like MBSR and MBCT to local adaptations, and pooled them into a single estimate. The headline was encouraging: a standardized mean difference somewhere between .46 and .73, a moderate-to-large effect, the kind of number that has become “mindfulness works” by the time it reaches a textbook or a clinician.

But look at the question that average answers — “the effect of mindfulness therapy” — and you will recognize the grammar from the previous section. It treats “mindfulness therapy,” unconditioned, as one homogeneous thing, an effect that exists and the meta-analysis merely measures. The studies it pooled were not one thing. Mindfulness for chronic pain in an American clinic and mindfulness for depression in a Chinese university are no more the same treatment of the same disorder in the same people than a newborn and an adult are the same height. The pool is heterogeneous by construction — across disorders, populations, and therapies at once. Asking for its single average effect is asking for the average height of everyone on Earth.

Hidden Moderators in Plain Sight

A meta-analyst can fairly say that I have described only half the job. Meta-analysts do not just compute an average; they also look for moderators — study features that predict when the effect is larger or smaller. Culture, dosage, type of control group, severity of the disorder: code each study on these characteristics, then test whether they track the effect sizes. This is the right instinct. If effects vary, find out what they vary with. Sometimes this works. But in psychology it often does not, and the reason is partly built into the way the search works.

A moderator analysis can only find variation associated with variables that were actually coded. You choose the study characteristics, code the studies on them, and ask which ones predict the results. That can reveal an explanation only if two things are true: you thought to measure it, and enough studies differ on it for a pattern to emerge. When the real source of variation is something no one thought to code — an unusual outcome measure, a quality problem, a researcher who strongly favored a particular result — the moderator analysis may come back empty.

Then there is variation that does not correspond to any broad study characteristic at all. Imagine making a smoothie with a dozen different fruits. Suppose it tastes off because of a single rotten blueberry — one study in forty with a broken measure, a p-hacked result, or invented data. There may be no useful moderator for that. “Rotten” is not a dimension along which the studies vary; it is a fact about one study. A moderator is a column in a spreadsheet, while a single unusual study is a row. To understand that study, eventually you have to look at the row.

Psychologists already know the shape of this problem from their own statistical tools. Factor analysis looks for variation that is shared across several measures. A strong relationship between just two variables does not ordinarily define a broad factor and may be treated as something specific to that pair. Cluster analysis asks a different question. If two variables correlate at .9, they can form a tight cluster whether or not they belong to any broader dimension. Moderator analysis resembles the factor approach: it looks for systematic variation along dimensions shared by multiple studies. It is less useful for a small pocket of studies that resemble one another for some idiosyncratic reason, and still less useful for a single unusual study. Those patterns become visible only when we stop looking exclusively at columns and start looking at rows.

In a meta-analysis of treatment effectiveness, however, the rotten blueberries are not the only studies we should be looking for. We also want the opposite — studies that provide especially strong and trustworthy evidence that the treatment works. But finding them is harder than sorting a forest plot by observed effect size. Large effects from small studies are especially vulnerable to sampling error, and extreme estimates are often extreme partly because of luck. Rank studies by the effects you happen to observe and you risk promoting the flukes.

This is where z-curve3 can help. It uses information from the distribution of results to shrink noisy study estimates toward more plausible values, correcting for regression to the mean and selective reporting. From the adjusted estimate it can compute a minimum effect size: a conservative lower-bound estimate of how large the effect could reasonably be after sampling error and uncertainty are taken into account. That makes a different kind of claim from the pooled average. The pooled mean asks for the center of the entire collection. The minimum effect size asks what can be said conservatively about one particular study.

And this brings us back to the blueberry. Moderator analysis asks which characteristics explain differences across studies. The corrected forest plot asks a different question: which individual studies provide the strongest evidence after noisy estimates have been pulled back toward more plausible values? The figure shows those studies, along with an estimate of how likely each result is to reach significance again in an exact replication of the same size. These are the promising fruits for a tasty smoothie. You find them not by blending everything together, and not only by coding broad dimensions, but by looking at the studies one at a time.

Do Western Patients Benefit from Eastern Mindfulness Therapy?

The figure shows a forest plot of the studies with the strongest evidence, sorted by their minimum effect size, from a high of 1.56 down to .41. Each study is identified by its first author and year. Look at which studies produced strong evidence of effectiveness on their own. Names like Majid, Zemestani, Kaviani, Omidi, Bakhshani, Panahi, Zhang, Chien, and Wang are Asian names, and closer inspection of the articles confirms it: these were studies of Asian participants. The strong evidence in this literature comes, overwhelmingly, from Iran and China. The pattern was sitting in the 2018 data; it took a 2021 umbrella review to note, across this body of work, that effects tend to run larger in Asian studies (Goldberg et al., 2021).

Given these results, a meta-analysis that pools all studies tells us nothing about the effectiveness of mindfulness therapy in WEIRD or in non-WEIRD samples. The average is too high for the Western patient, whose studies cluster low, and too low for the Iranian and Chinese patient, whose studies cluster high.

It may seem laudable that the meta-analysis included non-WEIRD samples. But dropping them in the blender is what created the heterogeneity that makes the average useless in the first place. The pooled number tells us nothing about either population on its own — it is an average across both that describes neither. And the fix is not complicated. Before you average a set of studies, you owe one check: do their results scatter by luck alone? If the only thing separating the estimates is sampling error, the studies were plausibly measuring one effect, and the average means something. If they scatter by more than luck — if real differences remain after chance is accounted for — then they were never one thing, and no single number should be reported for all of them.

So do Western patients benefit from Eastern mindfulness therapy? This meta-analysis cannot say. This is not a verdict on mindfulness therapy. It is a verdict on a method. There is nothing weird about studying WEIRD samples, if the question is whether mindfulness therapy helps WEIRD patients. What is weird is to mix populations, discover that the effects vary, and then report the average as if it applied to all of them.

Conclusion

Science is a process. While there are universal criteria that distinguish science from other belief systems, the universal aspect of science is to question itself and to learn from mistakes. This process can take time. Meta-analysis emerged in the 1970s to make sense of inconclusive and sometimes conflicting results in a growing literature of empirical studies. Over time, rules for meta-analyses were formulated. Nowadays, meta-analyses are often considered to be the gold standard to make sense of original studies and meta-analyses are highly cited as authoritative sources to make claims like “Mindfulness therapy works.”

Initial meta-analysis often assumed a single effect size. Over time, methods were developed to examine and quantify heterogeneity in population effect sizes. However, meta-analysts are still trying to figure out how to report heterogeneity and what to with it. This essay points out that heterogeneity in effect sizes cannot be ignored. Studies should be combined to reduce sampling error, but not to hide true variation across populations.

More broadly, psychologists need to become more comfortable to study specific populations rather than claiming that their study tests a universal hypothesis and then apologize for the fact that they studied only US Americans or another WEIRD population. Studies that do want to make universal claims (e.g., Ekman’s research on facial expression) do require cross-cultural data, but not all studies have to test universal hypotheses.

Further Readings

  • Ghai, S. (2021). “It’s time to reimagine sample diversity and retire the WEIRD dichotomy.” Nature Human Behaviour. This is probably the cleanest paper for your purpose. Ghai argues that dividing the world into WEIRD versus non-WEIRD collapses enormous heterogeneity into a binary classification. A sample from India, Nigeria, Chile, and rural China does not become meaningfully similar simply because all are “non-WEIRD.”
    Nature Human Behaviour article
  • Clancy, K. B. H., & Davis, J. L. (2019). “Soylent Is People, and WEIRD Is White: Biological Anthropology, Whiteness, and the Limits of the WEIRD.” Annual Review of Anthropology. This is a deeper conceptual critique. They argue that the individual components of WEIRD are poorly operationalized and that treating inhabitants of “WEIRD societies” as homogeneous erases substantial differences within those societies. Their broader argument is that the label can obscure the actual dimensions researchers need to measure.
    Annual Review article
  • Muthukrishna et al. (2020). “Beyond Western, Educated, Industrial, Rich, and Democratic (WEIRD) Psychology: Measuring and Mapping Scales of Cultural and Psychological Distance.” Psychological Science. This comes partly from the same intellectual tradition as the original WEIRD paper, but it implicitly identifies a major problem with the acronym: cultural variation is better conceived as multidimensional and continuous rather than as membership in two groups. They develop measures of psychological/cultural distance instead.
    Paper information and full-text links
  • Schimmelpfennig et al. (2024). “Methodological concerns underlying a lack of evidence for cultural heterogeneity in the replication of psychological effects.” Communications Psychology. This paper includes Henrich, Heine, and Norenzayan themselves. It explicitly warns against turning the letters of WEIRD into an empirical “WEIRDness” scale. Their point is important: WEIRD was originally a mnemonic/consciousness-raising device, not a theory of which cultural dimensions cause psychological variation. They criticize binary coding and mechanically decomposing countries according to the five letters because this produces classifications with poor theoretical and face validity. Open-access article
  • Jeffrey Sherman’s “There Is Nothing WEIRD About Basic Research: The Critical Role of Convenience Samples in Psychological Science” in American Psychologist (published online 2024; print 2025). Sherman accepts that psychology has a diversity problem, but challenges the inference that every study therefore requires culturally representative or highly diverse sampling. His argument is that the relevant question is what population a claim is intended to generalize to and what moderators the theory predicts. Convenience sampling can be entirely appropriate for basic research. He also stresses that “WEIRD sample” and “convenience sample” are not the same methodological problem.
  • Open manuscript copy

Credibility in Economics: A Reanalysis of Large-Scale Meta-Research

Askarov, Z., Doucouliagos, A., Doucouliagos, H., & Stanley, T. D. (2024). Selective and (mis)leading economics journals: Meta-research evidence. Journal of Economic Surveys, 38(5), 1567–1592. https://doi.org/10.1111/joes.12598

Abstract

Askarov, Doucouliagos, Doucouliagos, and Stanley (2024) analyzed statistical power and excess statistical significance in a large collection of economics meta-analyses and concluded that much of the evidence reported in leading economics journals is potentially misleading. We used their open data to conduct a z-curve analysis to examine the credibility of economics using a different statistical model. Z-curve has several advantages over the power-analysis and Test of Excess Significance (TES) approach used by Askarov et al. First, it does not assume that all studies within a meta-analysis share a single population effect size. Instead, it models heterogeneity with a mixture model. Second, z-curve models selection for statistical significance and uses the fitted distribution of significant results to estimate the discovery rate that would be expected in the absence of selection. The discrepancy between the observed and expected discovery rates therefore provides a direct measure of selection bias. In contrast, TES does not explicitly model how selection distorts the distribution of observed effect sizes when estimating expected significance. Its UWLS estimator gives greater weight to more precise estimates, which typically come from larger samples. If smaller, less precise studies report inflated effect sizes, the weighted mean will be pulled toward the smaller effects observed in more precise studies, thereby reducing the estimated power assigned to the smaller studies. This weighting can reduce small-study bias, but it does not necessarily eliminate selection bias. Moreover, if true effect sizes systematically differ with study size, the same weighting can itself produce a biased estimate of the average effect. Third, z-curve distinguishes between overall power (the Expected Discovery Rate, EDR) and power conditioned on significance (the Expected Replication Rate, ERR). With heterogeneous data, the average power of significant results can be much higher than overall power. Finally, z-curve uses the EDR to obtain an upper bound on the false discovery rate using a formula developed by Sorić (1989).

First, Askarov et al.’s estimate-level mean power and the z-curve EDR are surprisingly similar, approximately 27% and 28%, respectively. A discovery rate of this magnitude implies a maximum false discovery rate of approximately 14%. Second, the expected replication rate of statistically significant results is approximately 70%, showing that the power of selected significant results is substantially higher than overall power. These estimates are similar to estimates obtained for randomized clinical trials in medicine and do not support pessimistic interpretations of this database based solely on its low median power. Low overall power is primarily a problem for discovery: true effects are less likely to reach significance, creating the potential for false negatives. Importantly, Askarov et al.’s own database shows that many nonsignificant estimates are nevertheless reported and incorporated into meta-analyses, where evidence can be aggregated to increase precision and statistical power. Thus, low power of individual studies does not by itself imply low credibility of the resulting literature.

Introduction

Concerns about the credibility of science are no longer purely academic. Scientific evidence informs consequential decisions about health, climate, and economic policy, making the credibility of published research important for both policymakers and the public. Yet academic incentives can undermine credibility. Researchers are rewarded for novel and statistically significant findings, whereas replications and corrections receive less attention. As a result, false positive findings may enter the literature and persist even when later evidence fails to support them.

Concerns about scientific credibility intensified after Ioannidis (2005) argued that most published research findings are false. Although influential, this claim was largely theoretical rather than based on an empirical estimate of false discoveries across science. For most significant results to be false positives, researchers must test many false hypotheses and have relatively low power to detect true effects. For example, if only 10% of tested hypotheses are true, statistical power is 50%, and the Type I error rate is 5%, then 5% of the true hypotheses and 4.5% of the false hypotheses will produce significant results. Consequently, nearly half of all significant results, 4.5/(4.5 + 5) = 47%, would be false discoveries.

Empirical investigations of scientific credibility have produced a less pessimistic but highly variable picture. Button et al. (2013) documented very low statistical power in neuroscience, with median power estimates across meta-analyses ranging from approximately 8% to 31%. In contrast, Jager and Leek (2014) analyzed reported p-values in major medical journals and estimated that only 14% of significant results were false discoveries. Direct replication projects introduced yet another measure of credibility. The Open Science Collaboration (2015) found that only 36% of psychology findings produced a significant result in the same direction in a replication, whereas Camerer et al. (2016) obtained a replication rate of 61% for laboratory experiments in economics. A much larger recent investigation of the social and behavioural sciences found that approximately half of tested claims replicated.

Concerns about credibility have also become prominent in economics. Large meta-research projects have documented selection for statistical significance and low statistical power. Most recently, Askarov et al. (2024) analyzed 368 meta-analyses containing 167,753 estimates, including 22,281 estimates published in 31 leading economics journals. They emphasized that median power in the leading journals was only 7% and reported substantial excess statistical significance, leading them to question the credibility of much published economics research. At the same time, direct replication and robustness studies have produced more encouraging results. Camerer et al. (2016) replicated 61% of experimental findings, while a recent large-scale study found that 72% of significant economics and political-science estimates remained significant and in the same direction under alternative analyses. Thus, empirical assessments of economics range from very low estimates of statistical power to substantially higher estimates of replicability and robustness.

These quantities can differ substantially when statistical power is heterogeneous. Moreover, estimates from different methods depend on different assumptions about effect-size heterogeneity, selection for significance, and the proportion of true null hypotheses. Consequently, apparently conflicting estimates of scientific credibility need not actually contradict one another.

The present study addresses this problem using z-curve, a statistical model that estimates several credibility parameters within a single coherent framework. Z-curve models heterogeneity in statistical power with a mixture distribution and explicitly models selection for statistical significance. Its main estimands are the EDR and the ERR. The EDR can be compared with the Observed Discovery Rate (ODR), the percentage of significant results, to assess and quantify selection for statistical significance. Furthermore, the EDR can be used to estimate the maximum False Discovery Risk (FDR) using a formula developed by Sorić (1989). We use the term risk rather than rate because the actual false discovery rate cannot be identified from the observed test statistics alone without knowing which tested null hypotheses are true.

Sorić’s formula shows that the relationship between EDR and maximum FDR is nonlinear. For example, an EDR of 20% implies a maximum FDR of approximately 21% at α=.05. Thus, even low mean discovery probabilities do not imply that most significant results are false positives.

Data

The Askarov et al. dataset combines 368 economics-related meta-analyses covering a broad range of research areas. The meta-analyses were identified through bibliographic databases, publisher websites, specialist journals, and searches of work by known meta-analysts; the search ended on July 31, 2021. When data were not publicly available, the authors contacted the original meta-analysts and obtained data from 74% of those contacted. To be included, a meta-analysis had to contain at least five primary studies and report both effect-size estimates and their standard errors. When multiple meta-analyses examined the same research area, the most recent and comprehensive one was selected. The final dataset contains 167,753 estimates, including 22,281 estimates published in 31 leading general-interest and field economics journals. The authors emphasize that the dataset is not necessarily representative of all empirical economics research, but rather of research areas that have been subjected to meta-analysis.

The database also contains identifiers for the original primary studies, making it possible to account for dependence among multiple estimates reported by the same study. The 167,753 estimates represent approximately 15,000 primary-study clusters.

Results

The most important estimate is the Expected Discovery Rate (EDR) of 27%. This estimate means that an unbiased sample of tests drawn from the same underlying population is expected to contain approximately 27% significant results. This estimate is surprisingly close to Askarov et al.’s estimate-level mean power of approximately 27%.

The two quantities are conceptually similar but not identical. Askarov et al. calculate directional power: significance is counted only in the direction of the estimated meta-analytic effect. Z-curve’s EDR uses two-sided statistical significance. Consequently, the null baseline for Askarov et al.’s directional calculation is 2.5%, rather than the conventional two-sided Type I error rate of 5%. The numerical difference between directional and two-sided power becomes very small as power increases, however, and does not explain the close agreement between the aggregate estimates.

The similarity of the mean estimates is particularly informative because the mean, rather than the median, determines the expected proportion of significant results.

For the full database, approximately 51% of reported estimates are significant, whereas z-curve estimates an EDR of 27%, a difference of approximately 24 percentage points. Askarov et al.’s estimate-level mean power for the full database is also approximately 27%, implying a very similar aggregate discrepancy between observed and expected significance. This numerical agreement should not be interpreted as validation of the two methods. Askarov et al. calculate power from a common meta-analytic effect within each research area, whereas z-curve estimates a heterogeneous distribution of noncentrality parameters. The two approaches can therefore produce very different results in individual heterogeneous meta-analyses even when their aggregate averages happen to agree.

It is unconventional to refer to the difference between observed and expected significance as a “rate of false positives.” The term false positive normally refers to a statistically significant result that incorrectly rejects a true null hypothesis. Excess significance does not establish that the excess results are false rejections of H0​. They may instead reflect inflated estimates of real effects caused by selective reporting or specification searching. Thus, Askarov et al.’s excess-significance measure should not be interpreted as an estimate of the proportion of significant findings that are false discoveries.

In contrast, z-curve uses the EDR to estimate an upper bound on the proportion of significant results that could be false discoveries. Following Sorić (1989),FDRmax​=(EDR1​−1)1−αα​.

With an EDR of 27%, the maximum FDR is approximately 14%. Allowing for sampling uncertainty in the EDR raises the upper confidence limit to approximately 19%. Thus, the results imply that no more than roughly one in five significant results could be false discoveries within the assumptions of the model. The actual FDR may be considerably lower. The Sorić bound is obtained under the extreme assumption that true alternatives are detected with perfect power; when power against true alternatives is lower, fewer of the observed significant findings can be attributed to true null hypotheses.

The most dramatic difference between Askarov et al.’s interpretation and the z-curve results concerns their emphasis on median power. Askarov et al. highlight median power of only 7% in leading economics journals and note that this value is close to the conventional 5% significance criterion. This comparison is misleading for two reasons.

First, their power calculation is directional. Under a true null hypothesis, their formula produces a probability of 2.5%, not 5%. Thus, a directional power estimate of 7% should not be compared directly with the two-sided Type I error rate of 5%. This distinction has little impact once power becomes moderate, but it matters for interpreting values very close to the null.

Second, and more importantly, median power is not the quantity that predicts how many significant results a literature should produce. The mean probability of significance does. Their own estimate-level mean power is approximately 27%, nearly four times their headline median of 7% and remarkably close to the z-curve EDR.

The distinction also matters for credibility. A low discovery probability across all tests implies that many results will be nonsignificant. This is a serious problem when nonsignificant findings are suppressed, because selective reporting will exaggerate the apparent success of the literature. But low discovery probability does not imply that significant findings themselves have similarly low replicability.

Z-curve estimates the Expected Replication Rate of significant results at 69%. Thus, although the EDR for all tests is only 27%, results that passed the significance threshold are estimated to have substantially higher power. The distinction follows directly from selection: results with higher underlying power are more likely to become significant and therefore are overrepresented among significant findings.

The ERR also includes any true null results that happened to become significant. At the maximum-FDR point estimate of 14%, the implied same-direction replication probability among the remaining true-positive results would be approximately 80%. This calculation should not be interpreted as a separate estimate of the true-positive power because the 14% FDR is itself an upper bound. It simply illustrates that low overall discovery probability can coexist with much higher replicability among significant results that reflect genuine effects.

In short, evaluations of credibility need to distinguish among several quantities: the probability of significance across all tests, the probability of significance among true alternatives, the replicability of results selected for significance, and the probability that a significant result is a false discovery. Median discovery probability provides little information about the latter two quantities.

Askarov et al.’s finding of low median power therefore does not by itself imply that economics research lacks credibility. Their own mean-power estimate and the z-curve EDR both suggest an underlying discovery probability of approximately 27%, while z-curve estimates an ERR of approximately 69% and a maximum FDR of approximately 14%. These results indicate substantial selection for statistical significance and considerable room for improvement, but they do not support the conclusion that the low median power of individual estimates, by itself, raises serious doubts about the credibility of the meta-analyzed economics literature.

Conclusion

In conclusion, meta-scientists often point out that extraordinary claims require extraordinary evidence and that academic incentives can reward researchers for making strong claims from weak evidence. Meta-science is not immune to these pressures. The claim that an entire discipline conducts studies with a typical probability of only 7% of rejecting a false null hypothesis is remarkable, if true. However, closer examination shows that this headline figure is a median discovery probability and is not the quantity that predicts the expected number of significant results or the credibility of significant findings. Askarov et al.’s own mean estimate is approximately 27%, closely matching the z-curve EDR, while z-curve estimates substantially higher replicability among significant results and a relatively modest upper bound on the false discovery rate. Thus, the evidence supports concerns about selective reporting and low discovery rates, but it does not support the much stronger conclusion that the low median power estimate by itself raises serious doubts about the credibility of economics research.

Primed for Equivocation

“Priming exercise” and psychological priming share a label, not necessarily a mechanism. Deliberate preparation for a known future competition provides no obvious need for unconscious goal activation, and the article presents no evidence that such activation actually occurs. The shared terminology leaves the proposed explanation primed for equivocation.

Holmberg and Kelly (2026) provide a useful critique of physiological explanations for “priming exercise,” but their attempt to connect this literature to psychological priming introduces a much less plausible mechanism. The connection appears to arise largely because the two literatures happen to use the same word. That shared terminology leaves the argument primed for equivocation.

In sports physiology, “priming exercise” refers to a deliberately performed bout of exercise intended to improve performance later that day. In cognitive and social psychology, priming refers to prior exposure to a stimulus that subsequently alters processing or behavior, sometimes without awareness or conscious intention. These are fundamentally different uses of the term. The fact that both involve something occurring before something else does not imply that they share a psychological mechanism.

This distinction becomes especially important when the authors invoke nonconscious goal priming. They suggest that exercise might activate performance goals and related behavioral representations that persist until later testing. They cite classic social-psychological priming research, including Bargh and colleagues, to support the possibility that goals can be activated without awareness and subsequently influence behavior.

But the proposed mechanism is poorly matched to the phenomenon being explained. An athlete does not ordinarily encounter a “priming exercise” incidentally. The athlete performs it because a competition or performance test is coming later. The later performance goal is therefore likely to have been activated before the exercise begins:

competition later → intention to prepare → priming exercise.

The goal is not plausibly dormant until the exercise somehow activates it unconsciously. Indeed, the goal is probably one of the reasons the athlete performs the exercise in the first place. During the exercise the athlete may also consciously think about the competition, technique, pacing, readiness, or expected benefits. Under those circumstances, invoking nonconscious goal activation is not merely unnecessary; the proposed causal sequence is almost backwards.

The authors’ own examples illustrate the problem. They discuss athletes rehearsing particular pacing strategies, using metronomes, receiving verbal cues, believing that squatting will improve subsequent jumping, and developing confidence or expectations about later performance. These are readily understood as deliberate preparation, task practice, expectancy, motivation, or attentional effects. None requires a nonconscious priming mechanism.

The scientific evidence offered for the nonconscious account is also weak. The article provides no exercise experiment demonstrating that a priming-exercise bout activates a previously inactive goal outside awareness and that this activation subsequently causes improved performance hours later. In fact, the authors repeatedly acknowledge that these possibilities “may” occur, are “hypothesized,” or “have yet to be directly examined.” Thus, the proposed mechanism is not an empirical finding from the exercise literature.

Instead, support is imported from a different literature on behavioral and goal priming. That literature is itself scientifically controversial, and citing classic demonstrations does not establish that the same mechanism operates in an entirely different situation involving intentional athletic preparation. The conceptual inference appears to be:

psychology calls something “priming”

  • exercise science calls something “priming”
    → psychological priming may explain exercise priming.

But identical terminology is not evidence of mechanistic equivalence.

The distinction matters because the paper already identifies much more plausible explanations. Task-specific practice, motor learning, expectancy, researcher effects, motivation, and explicit performance preparation could all produce later performance changes. These mechanisms fit the actual structure of the situation: athletes know that performance is coming and intentionally prepare for it. The nonconscious goal-priming hypothesis adds an unnecessary and poorly supported causal layer.

The problem can therefore be summarized simply: an implausible mechanism is invoked to explain effects that have not been shown to require that mechanism, and the empirical justification comes largely from a separate literature that happens to use the same word.

Methodological Problems in Claims About “Conscious” and “Preconscious” Influencer Effects

Mir, I. A. (2026). Influencer’s physical attractiveness and content aesthetics: Conscious and preconscious determinants of fashion-branded content engagement on Instagram. Journal of Creative Communications, 21(2), 201–219. https://doi.org/10.1177/09732586241288672

“The authors invoke an unproven perception–behavior mechanism to explain causation that was never observed. A direct path in a cross-sectional SEM establishes neither causation nor preconscious processing.”

Mir (2026) examines whether fashion influencers’ physical attractiveness and the aesthetics of their branded content predict followers’ engagement on Instagram. The study uses survey responses from 300 followers of 15 fashion influencers in Pakistan and analyzes the proposed relationships with structural equation modeling and mediation analyses. The main empirical finding is straightforward: followers who rate influencers and their content more positively also report more favorable attitudes and greater engagement. The methodological problem is that the article draws causal and psychological-process conclusions that the design cannot support.

The most serious problem is that all variables were measured in a single cross-sectional self-report survey. Participants simultaneously reported how attractive they considered the influencer, how aesthetically pleasing they considered the content, their attitude toward that content, and how often they viewed, liked, commented on, and shared it. Nothing was manipulated, and there was no temporal ordering of the variables. Nevertheless, the article repeatedly describes attractiveness and aesthetics as factors that “cause,” “trigger,” “stimulate,” or “activate” engagement. Those causal statements do not follow from the design.

For example, the proposed model assumes

attractiveness → attitude → engagement.

But the same covariance pattern is compatible with numerous alternatives. Followers who engage frequently with an influencer may develop more positive attitudes and subsequently rate that influencer as more attractive. A general liking or identification with the influencer could simultaneously increase attractiveness ratings, content-aesthetic ratings, attitudes, and engagement. The structural equation model cannot distinguish among these explanations.

The sampling procedure makes this problem particularly important. Participants had to have followed one of the selected influencers for more than six months. Thus, the sample is already conditioned on sustained interest in the influencer. People who disliked the influencer, found the content unattractive, or disengaged from it are systematically less likely to appear in the sample. The resulting correlations describe differences among an already selected group of followers; they provide weak evidence for the article’s practical recommendation that firms should hire physically attractive influencers because attractiveness causes engagement.

The article’s central methodological error is even more fundamental. It claims to distinguish a “conscious” route from a “preconscious” route. The mediated path

attractiveness/aesthetics → attitude → engagement

is interpreted as conscious influence, whereas a remaining direct path from attractiveness or aesthetics to engagement is interpreted as evidence of a preconscious perception–behavior process.

A direct regression coefficient is not a measure of unconscious processing.

If attractiveness predicts engagement after statistical adjustment for an explicit attitude measure, this merely shows residual covariance between those variables. That residual association could reflect measurement error in attitude, omitted mediators, stable preferences, common response tendencies, reverse causation, or numerous other processes. Nothing in the study measures awareness, intention, automaticity, processing speed, or participants’ ability to report the causes of their behavior. Consequently, the data provide no evidence that engagement was “preconscious” or unintentional.

For the same reason, the indirect path through an explicit attitude measure does not establish a conscious causal mechanism. Participants consciously completed the attitude questionnaire, but that does not mean the psychological process producing their engagement operated consciously. Statistical mediation and conscious psychological mediation are different concepts.

This problem is especially consequential because the “preconscious” interpretation is one of the article’s principal theoretical contributions. The authors explicitly invoke the perception–behavior literature, including Bargh et al. (1996), to justify the claim that a significant direct path demonstrates automatic behavior. But the current research contains none of the experimental procedures that would be required to test an automatic perception–behavior effect. The SEM therefore cannot adjudicate between conscious and unconscious processes.

The measures themselves also create substantial interpretive problems. On page 209, physical attractiveness is measured with “stylish,” “good looking,” “sexy,” and “elegant.” Content aesthetics is measured with “striking,” “wonderful,” “fascinating,” and “lovely,” while attitude is measured with “pleasant,” “good,” “likeable,” and “my favourite.” These constructs are conceptually and evaluatively intertwined. “Wonderful” and “lovely,” for example, are not narrowly aesthetic judgments, while “stylish” and “elegant” are not purely measures of physical attractiveness. Much of the model may therefore reflect a broad positive-evaluation factor rather than distinct psychological constructs connected by causal pathways.

The observed correlations are consistent with this concern. Physical attractiveness correlates .64 with content aesthetics and .64 with attitude; content aesthetics correlates .66 with attitude and .68 with reported content consumption. Demonstrating discriminant validity with the Fornell–Larcker criterion does not eliminate the possibility that halo effects or general liking strongly influence all of these ratings.

The attempt to dismiss common-method bias is also inadequate. All predictors, mediator variables, and outcomes were obtained from the same respondent at the same time using similar rating formats. The authors test whether several sets of items can be represented by a single latent factor and conclude that poor single-factor fit shows that common-method bias is not important. That conclusion does not follow. Common-method variance does not require every item to load on one factor. Several distinguishable constructs can coexist while correlations among them are inflated by shared method, evaluative consistency, acquiescence, or halo effects.

The proposed “snowball effect” suffers from the same causal problem. The authors find that self-reported consumption behaviors—viewing, reading comments, and liking—predict contribution behaviors such as commenting and sharing, and conclude that consumption gradually causes users to progress toward more active participation. Yet consumption and contribution were measured simultaneously. Someone who frequently comments and shares influencer content almost necessarily also views and consumes that content. A positive cross-sectional association therefore does not demonstrate a temporal progression from passive to active engagement. Testing a snowball process would require longitudinal evidence showing that earlier consumption predicts subsequent increases in contribution.

There may also be an unmodeled dependence problem. The 300 respondents followed one of only 15 macro-influencers. Followers of the same influencer are not necessarily independent observations. Influencers may differ systematically in appearance, production quality, follower demographics, posting frequency, and baseline engagement. Those influencer-level characteristics could generate correlations among respondent ratings. The reported SEM appears to treat all 300 followers as independent rather than accounting for clustering by influencer.

Another reporting issue concerns the engagement scale. The response categories are described as 1 = “very often,” 2 = “often,” 3 = “sometimes,” and 4 = “never.” Thus, larger numerical values indicate less engagement. Yet positive path coefficients are consistently interpreted as greater attractiveness and aesthetics producing greater engagement. The article does not clearly state in the reported method that these scores were reverse-coded. If they were reversed before analysis, that transformation should have been explicitly documented. If they were not, the substantive interpretation of the coefficients would be reversed.

Finally, the statistical success of the model should not be confused with strong evidence for the hypotheses. All five proposed hypotheses are supported, including the weakest direct attractiveness effect, β = .12, t = 2.19. The study was not preregistered, and the sample-size justification consists largely of the statement that N = 300 is sufficient for purposive sampling and structural equation modeling rather than an a priori power analysis tied to the focal effects. Good model-fit indices demonstrate that a specified covariance model can reproduce the observed covariance matrix; they do not establish that the arrows in Figure 2 represent the true causal processes.

The study therefore supports a much narrower conclusion than the article claims. Among existing long-term followers of fashion influencers, positive ratings of influencer attractiveness and content aesthetics are associated with positive attitudes and greater self-reported engagement. That descriptive association is plausible and potentially useful.

The study does not establish that physical attractiveness or content aesthetics cause engagement, that attitude mediates these effects causally, that any residual direct relationship reflects a preconscious perception–behavior mechanism, or that passive engagement develops over time into active contribution.

A suitable experimental design would manipulate influencer attractiveness and content aesthetics independently, randomly assign participants to conditions, measure actual engagement behavior, and include measures capable of testing awareness or automaticity. A longitudinal design would be required to test the proposed consumption-to-contribution “snowball” process. Without such evidence, the article’s strongest psychological claims are interpretations imposed on cross-sectional correlations rather than findings produced by the research design.

Auditing a Poor Audit of Elderly Priming

Costa, T. (2026). The Bayesian audit: Evaluating the proportionality of scientific claims to evidence—a case study on social priming and walking speed. Frontiers in Psychology, 17, 1799078. DOI: 10.3389/fpsyg.2026.1799078


This article applies Bayes’ theorem to one conveniently selected t value and calls the result a ‘Bayesian audit.’ Most readers may simply ignore it because it appeared in Frontiers in Psychology. Those who want a more substantive reason can point to this review.

Costa (2026) introduces a “Bayesian audit,” a six-step framework intended to evaluate whether the strength of scientific claims is proportional to the evidence supporting them. The idea is sensible. Statistical significance does not tell us how strongly we should believe a scientific claim, and surprising claims based on weak evidence deserve particularly careful scrutiny. Costa illustrates the proposed method with one of social psychology’s most famous findings: Bargh, Chen, and Burrows’s (1996) claim that priming college students with words related to old age caused them to walk more slowly afterward.

Unfortunately, the audit itself is problematic. It misrepresents important features of the original study, considers only a fraction of the available evidence, and reduces a question about the magnitude and robustness of an effect to a comparison between a null and an inadequately specified alternative hypothesis.

The first problem is surprisingly basic. Costa describes the original finding as based on a study with approximately t(28) = 2.0 and p ≈ .05. But Bargh et al. actually reported two elderly-priming experiments. In Experiment 2a, the comparison was t(28) = 2.86, p < .01. They then conducted Experiment 2b as a replication and again reported slower walking, t(28) = 2.16, p < .05. Costa appears to approximate the weaker second result while failing to mention the stronger first result or even that the original article contained two studies.

That is an odd starting point for an audit. If the purpose is to reconstruct how much evidence supported the claim in 1996, both original studies should be included.

Costa also incorrectly describes participants as being “subliminally exposed to words related to old age.” They were not. Participants consciously read words while completing a scrambled-sentence task. The claimed unconscious component was that participants supposedly did not realize that the elderly-related words subsequently affected their walking. Bargh et al. themselves explicitly distinguished this procedure from subliminal priming; Experiment 3 of their paper used genuinely subliminal presentation of faces.

This distinction matters because Costa uses the apparent implausibility of unconscious effects on motor behavior to motivate skeptical prior probabilities. One should at least characterize the causal claim correctly before assigning a prior to it.

The treatment of replication evidence is even more problematic.

Costa cites Doyen et al. (2012) and Harris et al. (2013) as subsequent replication attempts. Doyen et al. did replicate the elderly-walking paradigm. In a substantially larger study using automated measurement, they found essentially no priming effect. Their second experiment further suggested that experimenter expectations could influence the result.

Harris et al. (2013), however, did not replicate elderly priming at all. They attempted to replicate Bargh et al.’s 2001 high-performance goal-priming experiments, in which achievement words were supposed to improve performance on a cognitive task. Calling Harris et al. a replication of the elderly-walking finding is simply an error.

More importantly, why is a Bayesian audit conducted in 2026 based primarily on one t statistic from 1996?

There is now a substantial literature on behavioral priming. Dai et al. (2023), for example, meta-analyzed 351 studies and 862 effect sizes and concluded that behavioral priming effects could be detected across a large literature. Conversely, Mac Giolla et al. (2024) examined 70 close replication attempts of 49 social-priming findings. Ninety-four percent produced smaller effects than the originals, only 17% were significant in the predicted direction, and among 52 replications conducted without an original author, none was significant in the original direction; the pooled effect for those independent replications was essentially zero.

These sources do not necessarily settle the question. Meta-analyses themselves can be distorted by publication bias and other forms of selection. But that is precisely why an audit should examine them critically. An audit of a 30-year-old scientific claim should evaluate the accumulated evidence, not simply convert one selected original result into a Bayes factor.

There is an even more fundamental problem with the statistical question Costa asks.

Costa assigns prior probabilities of .05, .10, and .20 to the alternative hypothesis and combines these with an estimated Bayes factor of approximately 3. This yields posterior probabilities of .14, .25, and .43, respectively. The arithmetic is straightforward. The interpretation is not.

Why should the prior probability that the effect exists be .05 or .10?

Costa acknowledges that these values are illustrative rather than derived from an elicitation procedure. But these priors largely determine the conclusion that posterior belief remains low. Starting with a 5% probability and multiplying the prior odds by a Bayes factor of 3 inevitably produces a low posterior probability.

More importantly, what exactly is the hypothesis whose prior probability is 5%?

There is a major difference between these propositions:

elderly-related words have exactly zero effect on walking speed;

elderly-related words have some nonzero effect;

elderly-related words have a psychologically meaningful effect;

elderly priming produces effects of the magnitude originally reported;

automatic stereotype activation reliably produces consequential behavioral changes.

These are not the same hypothesis.

The scientifically interesting issue today is probably not whether the population effect is exactly zero. The effect could be d = .05 or d = .10. Such an effect would make the point null hypothesis technically false while providing little support for the dramatic theoretical interpretation of the original experiments.

This is why effect sizes matter. Bargh et al.’s original studies implied very large effects. Subsequent evidence raises the possibility that the true effect, if it exists at all, is much smaller. A useful audit therefore needs to ask how large the effect is and how precisely it has been estimated—not merely whether H0 or H1 receives the larger Bayes factor.

Costa’s procedure also conflates two different kinds of priors. One is the prior model probability: how likely H1 is relative to H0 before seeing the data. The other is the prior distribution over possible effect sizes within H1. A Bayes factor for a composite alternative necessarily depends on the latter. Yet the article emphasizes sensitivity to prior model probabilities while giving much less attention to the effect-size assumptions used to obtain BF₁₀ ≈ 3.

This creates another problem with Costa’s distinction between “evidence” and “belief.” He describes the Bayes factor as quantifying evidence supplied by the data and posterior probability as combining this evidence with prior belief. But a Bayes factor is not simply a property of the observed data. It depends on the statistical models being compared, including the distribution of effect sizes assumed under the alternative hypothesis.

There is also an internal inconsistency in the treatment of the replication evidence. Costa states that the later replication attempts yielded Bayes factors close to 1 and therefore had little evidential impact. That makes sense: a Bayes factor of 1 leaves prior odds unchanged. Yet the subsequent synthesis says that posterior belief “collapses under replication.” It cannot do both. Replications with BF ≈ 1 cannot cause posterior belief to collapse. To demonstrate such a decline, one would need Bayes factors favoring the null or another competing model and then accumulate this evidence formally.

Publication bias is another conspicuous omission. The article is motivated by the replication crisis and explicitly acknowledges that biased data limit the usefulness of evidential measures. Yet the actual Bayesian calculation treats the published Bargh result as though it were an observation selected independently of statistical significance.

That is unrealistic. A BF of 3 obtained from a randomly selected study and a BF of 3 obtained from a literature in which statistically significant and theoretically exciting findings were preferentially published do not have the same evidential implications. If selection contributed to the replication crisis, an audit of the original published evidence needs to take selection seriously.

Costa also describes the original study’s low statistical power as an additional reason for skepticism. Low power certainly matters because significant results from low-powered studies tend to exaggerate effect sizes, particularly in a selected literature. But sample size has already entered the likelihood used to compute the Bayes factor. Low power is therefore not independent evidence against the hypothesis. The additional concern arises from selection, analytic flexibility, measurement error, and effect-size inflation.

The deeper problem is that Costa reduces a scientific question to H0 versus H1 when several competing explanations exist. Doyen et al.’s work raised experimenter expectancy as one possible explanation. Other possibilities include a genuinely small priming effect, effects restricted to particular conditions or individuals, procedural artifacts, or some combination of these mechanisms. A Bayes factor contrasting an exact-zero model with a generic nonzero-effect model cannot determine which causal explanation is correct.

This is especially important because rejecting H0 would not establish Bargh’s theory. Even convincing evidence for a tiny difference in walking speed would not demonstrate that automatic stereotype activation generally controls overt behavior.

The Bayesian audit is therefore based on a reasonable principle but a poor demonstration. Scientific claims should indeed be proportional to evidence. The problem is that assessing proportionality requires accurately identifying the original evidence, considering the accumulated replication literature, evaluating publication bias, distinguishing statistical from substantive hypotheses, and estimating plausible effect sizes and their uncertainty.

Ironically, the elderly-priming case illustrates the weakness of Costa’s audit more effectively than it illustrates its strengths. A proper audit should not ask merely whether one selected t statistic changes the odds that an effect is exactly zero. It should ask what three decades of evidence tell us about the magnitude, robustness, boundary conditions, and causal interpretation of the phenomenon.

On those questions, uncertainty remains. There may be a small elderly-priming effect. The evidence does not establish that the effect is exactly zero. But neither does the accumulated evidence support taking the spectacular effects reported in 1996 at face value. The important scientific task is to estimate what effect remains after accounting for uncertainty and bias. That requires more than Bayes’ theorem applied to one conveniently chosen t value.

Old Evidence for a Fragile Priming Theory

Przybylinski, E. (2026). Whatever you say: Changing transference-based problem behavior with if–then plans. Self and Identity. Advance online publication. https://doi.org/10.1080/15298868.2026.2613846


Przybylinski (2026) reports two experiments examining whether implementation intentions can prevent problematic behaviors triggered by transference. The theoretical logic is straightforward: subtle resemblance to a significant other is assumed to activate that person’s representation automatically, which can then influence memory, goals, and behavior. An if–then plan is proposed to prevent the activated representation from guiding subsequent behavior.

The principal concern is the credibility of the evidence on which this argument rests.

The article treats automatic behavioral priming as a well-established foundation. For example, it cites Bargh et al. (1996), Bargh et al. (2001), Chartrand and Bargh (1999), and related studies as evidence that contextual cues can automatically trigger overt behavior without awareness or intention. Yet behavioral priming is precisely one of the areas most affected by the replication crisis. The article does not discuss this change in evidential status. Thus, evidence that was considered persuasive when these studies were conducted is largely presented in 2026 as though subsequent replication failures had not occurred.

This matters because the two experiments are themselves products of that earlier research era. The author explicitly states that they were conducted as dissertation research in the late 2000s, before preregistration became standard. Study 1 included only 60 participants, or 20 participants in each of the three strategy conditions, despite testing interaction hypotheses. Study 2 included 47 participants. The manuscript provides a power justification, but because the studies were not preregistered, it is unclear whether the reported power analysis reflects a prospectively specified design decision. The reported “post hoc power” provides little additional information.

The results are remarkably successful. The focal interactions involving behavioral readiness, memory, and overt submissive behavior are all statistically significant in the predicted direction. The reported effects are also very large, frequently exceeding d = 1 and reaching d = 1.69 in Study 1. Across the major focal tests, the success rate is effectively 100%, while average observed power based on the reported effects is roughly 80%.

A perfect success rate is not impossible when power is 80%, but it is more successful than expected. Schimmack (2012) emphasized that unusually high success rates relative to estimated power can indicate that published effect sizes and success rates should not be taken at face value. Here, the outcomes within each experiment are dependent, so a simple excess-success or incredibility calculation would not be appropriate. With only two studies, there is no statistical smoking gun. Nevertheless, the combination of small samples, large effects, multiple significant focal outcomes, and absence of preregistration warrants substantial caution.

The historical timing makes this concern more important. These experiments were conducted approximately 15 years before their publication. During that interval, psychology experienced a replication crisis that directly challenged the credibility of the behavioral-priming literature on which the article relies. Yet the 2026 article does not supplement the old experiments with a contemporary, adequately powered, preregistered replication.

This is particularly striking because the central experiment is readily replicable. The author remains at the same institution where the original research was conducted, and the paradigm requires undergraduate participants rather than an unusually difficult population. A preregistered replication with a substantially larger sample could have provided highly informative evidence about whether the large effects observed in the original dissertation studies survive contemporary scrutiny.

The absence of such a replication changes how the evidence should be interpreted. The results are not invalid merely because they were collected before the replication crisis. But neither the reported effect sizes nor the perfect pattern of statistical success should be treated as reliable estimates of the underlying effects without independent replication.

A Speculative Theory of Replication Failures in Social Psychology: The Empty-Self Metaphor

Klein, J. W., & Swann, W. B., Jr. (2026). Social psychology’s empty-self metaphor and the replication crisis. Perspectives on Psychological Science, 21(2), 138–153. https://doi.org/10.1177/17456916251401849.

Klein and Swann (2026) offer an interesting diagnosis of the replication crisis, but the evidence does not support the strength of their theoretical interpretation.

Their central empirical observation is striking. They coded 41 hypotheses from Many Labs 1 and 2 as either consistent with an “empty-self” metaphor or not. None of the nine “empty-self” hypotheses replicated, whereas 26 of 32 “not-empty-self” hypotheses replicated, a difference of 81 percentage points. Given the small number of empty-self studies, however, the estimate is much less precise than the point estimate suggests; an approximate 95% confidence interval for the difference is about 47 to 91 percentage points. The authors appropriately describe the result as preliminary.

The more serious problem is interpretation. The analysis is correlational. Studies were not randomly assigned to use an “empty-self” theory, and the authors did not code plausible confounding variables that could themselves predict replication success. These include the sample size and statistical power of the original study, the strength of the situational manipulation, whether the manipulation was consciously perceived, the causal proximity between manipulation and outcome, and the prior plausibility of the predicted effect. Their own coding examples illustrate the problem. A gray background affecting support for austerity is coded as “empty self,” whereas paying for a workshop affecting attendance is coded as “not empty self.” The latter is still a situational effect. What differs most obviously is that payment is a strong and behaviorally relevant manipulation, whereas background color is a weak and remote one.

Thus, the empirical result may show that studies proposing large effects of weak, incidental situational manipulations replicate poorly. That is interesting, but it is not the same as showing that studies fail because they neglect an enduring self.

The theoretical explanation is also speculative and at times internally strained. Klein and Swann sometimes treat replication failures as evidence that subtle situational manipulations have little or no effect. In the Many Labs studies, this inference can be justified when very large samples produce narrow confidence intervals around zero. But this cannot be generalized to all failed replications. In other literatures, including elderly priming, the available confidence intervals may still be compatible with small effects. Failure to obtain significance is not equivalent to demonstrating an effect of exactly zero.

At the same time, Klein and Swann suggest that subtle situational effects may depend on the person. They argue that people may respond only to cues to which they are “tuned” and that primes may work only when they connect to existing self-representations. That possibility is not new to priming theory. Priming effects have long been assumed to depend on whether participants possess the relevant stereotype or representation, and some priming studies have explicitly predicted interactions—for example, effects of religious primes that depend on participants’ religiosity.

But this moderation account has an important implication. If a prime affects people for whom it is relevant and has little effect on others, the population-average effect should normally be reduced, not eliminated. Unless one assumes theoretically unusual crossover interactions in which the prime produces effects in opposite directions for different people, sufficiently large studies should still detect a small average effect. For elderly priming, for example, it is easy to imagine that some participants might be more responsive to an elderly stereotype than others. It is much harder to explain why the same prime should make another substantial group walk faster.

This creates a useful empirical question that Klein and Swann do not examine. Their 41 hypotheses should be coded not only as “empty self” or “not empty self,” but also according to whether the original hypothesis predicted a main effect or a Person × Situation interaction. If nearly all of the “empty-self” studies in Many Labs tested simple main effects, that matters because the larger priming literature already contains many moderator and interaction hypotheses. The Many Labs sample would then not represent the full theoretical range of priming research. More broadly, the authors should distinguish studies proposing universal effects from studies explicitly predicting conditional effects.

A stronger analysis would therefore code each study independently for situation strength, awareness of the manipulation, personal relevance, causal distance between manipulation and outcome, original sample size and power, prior plausibility, and whether the prediction was a main effect or an interaction. Only then could one determine whether an “empty-self” construct predicts replication after plausible alternative explanations have been taken into account.

The irony is that Klein and Swann criticize social psychology for building theories on weak evidence, but their own “empty-self” explanation is itself a speculative theory supported by weakly diagnostic data. The empirical pattern—0% versus 81% replication—is interesting. The claim that this pattern is caused by neglect of an enduring self is not established.

On a 1-to-5 scale from speculative theory to theory explaining highly credible phenomena, I would rate the article about 2/5. The phenomenon to be explained is credible: some classes of social-psychological findings replicate poorly. The proposed explanation—that they fail because social psychology adopted an “empty-self” metaphor—remains largely speculative.

A Mostly Speculative Social Identity Theory of Digital Identity

“Theories are a dime a dozen” (Ed Diener, personal communication)

Bingley, W. J., Worthy, P., Wiles, J., & Haslam, S. A. (2026). A social identity theory of digital identity. Perspectives on Psychological Science, 21(4), 346–382. https://doi.org/10.1177/17456916261419813.



Bingley and colleagues (2026) propose a new “social digital identity theory” (SDIT) to explain how social identities operate across online, offline, and hybrid environments. The theory is unusually explicit: the authors formulate 20 propositions concerning identity salience, digital platforms, embodiment, well-being, group functioning, polarization, culture, and other outcomes.

On a 1-to-5 continuum from speculative theory to theories that explain highly credible empirical phenomena (Speculative versus Explanatory Theory: A Review of “Narrative Embodiment” – Replicability-Index), I would give SDIT a rating of 2/5 (see .

This is not a judgment that the theory is uninteresting or implausible. A rating of 2 means that much of the distinctive theory remains speculative and that the empirical findings used to motivate it have not been systematically evaluated for credibility.

There are good features. Bingley et al. clearly distinguish established ideas from novel hypotheses. They explicitly acknowledge that many of their propositions—particularly P4 through P11, P18, and P20—have not previously been tested and require empirical evaluation. Other propositions are extensions of the much older social-identity and self-categorization literature. Thus, they do not claim that all 20 propositions are established facts.

The problem is different. When published studies are cited as empirical support, the authors rarely ask how strong that evidence actually is. There is no systematic consideration of statistical power, independent replication, publication bias, or whether an apparently supportive literature consists mainly of selected significant findings.

Embodiment provides a good example

Proposition 5 states that social-identity salience is shaped by embodiment. To motivate this proposition, the authors cite studies suggesting that heat activates anger concepts, social rejection makes rooms feel colder, keeping secrets feels physically burdensome, fist clenching activates related concepts, and facial-muscle activation influences experience. They then extend this reasoning to virtual embodiment and the Proteus effect.

A particularly revealing example concerns elderly priming.

The authors cite a virtual-reality study by Reinhard et al. in which participants embodied an older or younger avatar. Participants who had embodied the older avatar subsequently walked more slowly during the first part of a walking test than participants who embodied the younger avatar, p = .033. However, these differences disappeared during the second half of the walk.

Bingley et al. describe this result as operating similarly to “classic priming studies in social psychology” and cite Bargh, Chen, and Burrows (1996).

That citation is problematic because Bargh et al.’s famous claim that elderly-related primes cause people to walk more slowly became one of the best-known examples of the replication problems in social psychology. Doyen et al. (2012) conducted a larger replication using automated measurement and failed to reproduce the original effect in their first experiment.

None of this replication history is mentioned.

The result is an evidential chain that looks stronger on paper than it really is:

reported embodiment effects → elderly-avatar study → classic elderly priming → support for an embodiment proposition.

There are several citations, but citations are not the same thing as strong evidence.

The elderly-avatar study may be interesting, and the possibility that virtual embodiment affects subsequent behavior certainly deserves further study. But a marginally significant result that is then linked to a famous but poorly replicated priming finding does not provide strong evidence for a broad theoretical proposition about embodiment and social identity.

The same caution applies to meta-analytic evidence. A meta-analysis can summarize a literature very precisely while still giving a misleading estimate if the underlying literature is affected by publication bias and selective reporting. The important question is therefore not simply whether a theory can cite studies—or even a meta-analysis—that support a proposition. The question is whether the phenomenon itself has been demonstrated with credible evidence.

Why 2 rather than 1?

SDIT deserves more than the lowest score because it is not free-form speculation. It builds on established theoretical traditions, makes explicit predictions, distinguishes some novel claims from older ones, and repeatedly identifies propositions that require future empirical tests.

But it does not deserve a middle or high score because much of the distinctive theory is not yet explaining firmly established empirical phenomena. Moreover, the embodiment section demonstrates that the authors sometimes treat published findings as evidence without critically assessing whether those findings survived the replication crisis.

Thus, the 2/5 rating means:

SDIT is a structured and testable theoretical proposal with some grounding in established research, but much of its distinctive content remains speculative, and the empirical literature used to motivate some propositions is treated too uncritically.

The citation of Bargh et al. (1996) is a particularly clear example. In 2026, a theory article should not invoke elderly priming as supportive evidence without informing readers that the original phenomenon itself remains empirically uncertain.

Individual Differences Do Not Salvage The Priming Train Wreck

Aytürk, E., & Saribay, S. A. (2026). Neglect of the individual as a neglected problem: The relevance of combined idiographic-nomothetic approaches for social psychology today. Self and Identity. https://doi.org/10.1080/15298868.2026.2697941

Aytürk and Saribay (2026) make a useful methodological argument in “Neglect of the individual as a neglected problem.” Social psychology typically averages across people even though the same situation may have different psychological meanings for different individuals. They advocate greater use of personalized stimuli and intensive repeated-measures designs to identify person-specific processes before generalizing across people.

They also suggest that this problem might help explain replication failures in social priming. For example, “elderly” might evoke a frail grandparent for one person and a vigorous retiree for another. Averaging across these individuals could obscure different or even opposing effects. This is a reasonable hypothesis. It is also useful that the authors explicitly acknowledge low statistical power, questionable research practices, and publication bias as established contributors to replication failures.

The problem is that they nevertheless cite Bargh et al. (1996) rather uncritically. The famous finding that activating the elderly stereotype makes people walk more slowly is treated as an example of a potentially heterogeneous priming effect, without informing readers that the original effect became a prominent example of psychology’s replication problems. Doyen et al. (2012), using a larger sample and automated measurement, failed to reproduce the original effect in their first experiment.

There were earlier attempts to explain the inconsistent findings by invoking individual differences. For example, Cesario et al. (2006) reported that elderly priming slowed participants with relatively positive attitudes toward elderly people but sped up those with negative attitudes. Other studies similarly reported moderator effects. But these small-sample studies do not provide strong evidence that a robust main effect was merely hidden by heterogeneity. Cesario et al.’s study, for example, had only about 67 participants for the relevant analysis and did not obtain a conventionally significant overall priming effect. The evidential burden therefore shifted to an interaction estimated from an even less powerful design.

This is important because interaction effects are generally more difficult to estimate precisely than main effects. Finding a significant personality moderator in a small study does not establish that a failed main effect was really caused by individual differences. The moderator itself needs adequate power and independent replication. A more detailed discussion of these purported replications of elderly priming is available in “Elderly Priming: Did It Ever Work?”

Here Aytürk and Saribay’s methodological proposal may actually provide a better way forward. Intensive repeated measurement can obtain much more information from each participant and therefore can study within-person Person × Situation effects with fewer participants than a conventional design may require. This comes at a cost: many observations per person are needed, and ecological or experience-sampling studies typically sacrifice some of the experimental control available in tightly controlled laboratory experiments. Moreover, any between-person moderator still ultimately depends on having enough individuals.

Thus, the proposal itself is worth pursuing. Low power means that the failed priming literature often provides lack of evidence rather than definitive evidence of absence. Small, conditional, person-specific priming effects remain possible.

But they remain hypotheses to be demonstrated. Individual heterogeneity can explain variation in a real effect; it cannot by itself establish that an unreliable effect is real. A 2026 article explicitly concerned with the replication crisis should therefore not cite Bargh et al. (1996) without acknowledging the replication failures that make the existence of the phenomenon itself uncertain.