Askarov, Z., Doucouliagos, A., Doucouliagos, H., & Stanley, T. D. (2024). Selective and (mis)leading economics journals: Meta-research evidence. Journal of Economic Surveys, 38(5), 1567–1592. https://doi.org/10.1111/joes.12598
Abstract
Askarov, Doucouliagos, Doucouliagos, and Stanley (2024) analyzed statistical power and excess statistical significance in a large collection of economics meta-analyses and concluded that much of the evidence reported in leading economics journals is potentially misleading. We used their open data to conduct a z-curve analysis to examine the credibility of economics using a different statistical model. Z-curve has several advantages over the power-analysis and Test of Excess Significance (TES) approach used by Askarov et al. First, it does not assume that all studies within a meta-analysis share a single population effect size. Instead, it models heterogeneity with a mixture model. Second, z-curve models selection for statistical significance and uses the fitted distribution of significant results to estimate the discovery rate that would be expected in the absence of selection. The discrepancy between the observed and expected discovery rates therefore provides a direct measure of selection bias. In contrast, TES does not explicitly model how selection distorts the distribution of observed effect sizes when estimating expected significance. Its UWLS estimator gives greater weight to more precise estimates, which typically come from larger samples. If smaller, less precise studies report inflated effect sizes, the weighted mean will be pulled toward the smaller effects observed in more precise studies, thereby reducing the estimated power assigned to the smaller studies. This weighting can reduce small-study bias, but it does not necessarily eliminate selection bias. Moreover, if true effect sizes systematically differ with study size, the same weighting can itself produce a biased estimate of the average effect. Third, z-curve distinguishes between overall power (the Expected Discovery Rate, EDR) and power conditioned on significance (the Expected Replication Rate, ERR). With heterogeneous data, the average power of significant results can be much higher than overall power. Finally, z-curve uses the EDR to obtain an upper bound on the false discovery rate using a formula developed by Sorić (1989).
First, Askarov et al.’s estimate-level mean power and the z-curve EDR are surprisingly similar, approximately 27% and 28%, respectively. A discovery rate of this magnitude implies a maximum false discovery rate of approximately 14%. Second, the expected replication rate of statistically significant results is approximately 70%, showing that the power of selected significant results is substantially higher than overall power. These estimates are similar to estimates obtained for randomized clinical trials in medicine and do not support pessimistic interpretations of this database based solely on its low median power. Low overall power is primarily a problem for discovery: true effects are less likely to reach significance, creating the potential for false negatives. Importantly, Askarov et al.’s own database shows that many nonsignificant estimates are nevertheless reported and incorporated into meta-analyses, where evidence can be aggregated to increase precision and statistical power. Thus, low power of individual studies does not by itself imply low credibility of the resulting literature.
Introduction
Concerns about the credibility of science are no longer purely academic. Scientific evidence informs consequential decisions about health, climate, and economic policy, making the credibility of published research important for both policymakers and the public. Yet academic incentives can undermine credibility. Researchers are rewarded for novel and statistically significant findings, whereas replications and corrections receive less attention. As a result, false positive findings may enter the literature and persist even when later evidence fails to support them.
Concerns about scientific credibility intensified after Ioannidis (2005) argued that most published research findings are false. Although influential, this claim was largely theoretical rather than based on an empirical estimate of false discoveries across science. For most significant results to be false positives, researchers must test many false hypotheses and have relatively low power to detect true effects. For example, if only 10% of tested hypotheses are true, statistical power is 50%, and the Type I error rate is 5%, then 5% of the true hypotheses and 4.5% of the false hypotheses will produce significant results. Consequently, nearly half of all significant results, 4.5/(4.5 + 5) = 47%, would be false discoveries.
Empirical investigations of scientific credibility have produced a less pessimistic but highly variable picture. Button et al. (2013) documented very low statistical power in neuroscience, with median power estimates across meta-analyses ranging from approximately 8% to 31%. In contrast, Jager and Leek (2014) analyzed reported p-values in major medical journals and estimated that only 14% of significant results were false discoveries. Direct replication projects introduced yet another measure of credibility. The Open Science Collaboration (2015) found that only 36% of psychology findings produced a significant result in the same direction in a replication, whereas Camerer et al. (2016) obtained a replication rate of 61% for laboratory experiments in economics. A much larger recent investigation of the social and behavioural sciences found that approximately half of tested claims replicated.
Concerns about credibility have also become prominent in economics. Large meta-research projects have documented selection for statistical significance and low statistical power. Most recently, Askarov et al. (2024) analyzed 368 meta-analyses containing 167,753 estimates, including 22,281 estimates published in 31 leading economics journals. They emphasized that median power in the leading journals was only 7% and reported substantial excess statistical significance, leading them to question the credibility of much published economics research. At the same time, direct replication and robustness studies have produced more encouraging results. Camerer et al. (2016) replicated 61% of experimental findings, while a recent large-scale study found that 72% of significant economics and political-science estimates remained significant and in the same direction under alternative analyses. Thus, empirical assessments of economics range from very low estimates of statistical power to substantially higher estimates of replicability and robustness.
These quantities can differ substantially when statistical power is heterogeneous. Moreover, estimates from different methods depend on different assumptions about effect-size heterogeneity, selection for significance, and the proportion of true null hypotheses. Consequently, apparently conflicting estimates of scientific credibility need not actually contradict one another.
The present study addresses this problem using z-curve, a statistical model that estimates several credibility parameters within a single coherent framework. Z-curve models heterogeneity in statistical power with a mixture distribution and explicitly models selection for statistical significance. Its main estimands are the EDR and the ERR. The EDR can be compared with the Observed Discovery Rate (ODR), the percentage of significant results, to assess and quantify selection for statistical significance. Furthermore, the EDR can be used to estimate the maximum False Discovery Risk (FDR) using a formula developed by Sorić (1989). We use the term risk rather than rate because the actual false discovery rate cannot be identified from the observed test statistics alone without knowing which tested null hypotheses are true.
Sorić’s formula shows that the relationship between EDR and maximum FDR is nonlinear. For example, an EDR of 20% implies a maximum FDR of approximately 21% at α=.05. Thus, even low mean discovery probabilities do not imply that most significant results are false positives.
Data
The Askarov et al. dataset combines 368 economics-related meta-analyses covering a broad range of research areas. The meta-analyses were identified through bibliographic databases, publisher websites, specialist journals, and searches of work by known meta-analysts; the search ended on July 31, 2021. When data were not publicly available, the authors contacted the original meta-analysts and obtained data from 74% of those contacted. To be included, a meta-analysis had to contain at least five primary studies and report both effect-size estimates and their standard errors. When multiple meta-analyses examined the same research area, the most recent and comprehensive one was selected. The final dataset contains 167,753 estimates, including 22,281 estimates published in 31 leading general-interest and field economics journals. The authors emphasize that the dataset is not necessarily representative of all empirical economics research, but rather of research areas that have been subjected to meta-analysis.
The database also contains identifiers for the original primary studies, making it possible to account for dependence among multiple estimates reported by the same study. The 167,753 estimates represent approximately 15,000 primary-study clusters.
Results

The most important estimate is the Expected Discovery Rate (EDR) of 27%. This estimate means that an unbiased sample of tests drawn from the same underlying population is expected to contain approximately 27% significant results. This estimate is surprisingly close to Askarov et al.’s estimate-level mean power of approximately 27%.
The two quantities are conceptually similar but not identical. Askarov et al. calculate directional power: significance is counted only in the direction of the estimated meta-analytic effect. Z-curve’s EDR uses two-sided statistical significance. Consequently, the null baseline for Askarov et al.’s directional calculation is 2.5%, rather than the conventional two-sided Type I error rate of 5%. The numerical difference between directional and two-sided power becomes very small as power increases, however, and does not explain the close agreement between the aggregate estimates.
The similarity of the mean estimates is particularly informative because the mean, rather than the median, determines the expected proportion of significant results.
For the full database, approximately 51% of reported estimates are significant, whereas z-curve estimates an EDR of 27%, a difference of approximately 24 percentage points. Askarov et al.’s estimate-level mean power for the full database is also approximately 27%, implying a very similar aggregate discrepancy between observed and expected significance. This numerical agreement should not be interpreted as validation of the two methods. Askarov et al. calculate power from a common meta-analytic effect within each research area, whereas z-curve estimates a heterogeneous distribution of noncentrality parameters. The two approaches can therefore produce very different results in individual heterogeneous meta-analyses even when their aggregate averages happen to agree.
It is unconventional to refer to the difference between observed and expected significance as a “rate of false positives.” The term false positive normally refers to a statistically significant result that incorrectly rejects a true null hypothesis. Excess significance does not establish that the excess results are false rejections of H0. They may instead reflect inflated estimates of real effects caused by selective reporting or specification searching. Thus, Askarov et al.’s excess-significance measure should not be interpreted as an estimate of the proportion of significant findings that are false discoveries.
In contrast, z-curve uses the EDR to estimate an upper bound on the proportion of significant results that could be false discoveries. Following Sorić (1989),FDRmax=(EDR1−1)1−αα.
With an EDR of 27%, the maximum FDR is approximately 14%. Allowing for sampling uncertainty in the EDR raises the upper confidence limit to approximately 19%. Thus, the results imply that no more than roughly one in five significant results could be false discoveries within the assumptions of the model. The actual FDR may be considerably lower. The Sorić bound is obtained under the extreme assumption that true alternatives are detected with perfect power; when power against true alternatives is lower, fewer of the observed significant findings can be attributed to true null hypotheses.
The most dramatic difference between Askarov et al.’s interpretation and the z-curve results concerns their emphasis on median power. Askarov et al. highlight median power of only 7% in leading economics journals and note that this value is close to the conventional 5% significance criterion. This comparison is misleading for two reasons.
First, their power calculation is directional. Under a true null hypothesis, their formula produces a probability of 2.5%, not 5%. Thus, a directional power estimate of 7% should not be compared directly with the two-sided Type I error rate of 5%. This distinction has little impact once power becomes moderate, but it matters for interpreting values very close to the null.
Second, and more importantly, median power is not the quantity that predicts how many significant results a literature should produce. The mean probability of significance does. Their own estimate-level mean power is approximately 27%, nearly four times their headline median of 7% and remarkably close to the z-curve EDR.
The distinction also matters for credibility. A low discovery probability across all tests implies that many results will be nonsignificant. This is a serious problem when nonsignificant findings are suppressed, because selective reporting will exaggerate the apparent success of the literature. But low discovery probability does not imply that significant findings themselves have similarly low replicability.
Z-curve estimates the Expected Replication Rate of significant results at 69%. Thus, although the EDR for all tests is only 27%, results that passed the significance threshold are estimated to have substantially higher power. The distinction follows directly from selection: results with higher underlying power are more likely to become significant and therefore are overrepresented among significant findings.
The ERR also includes any true null results that happened to become significant. At the maximum-FDR point estimate of 14%, the implied same-direction replication probability among the remaining true-positive results would be approximately 80%. This calculation should not be interpreted as a separate estimate of the true-positive power because the 14% FDR is itself an upper bound. It simply illustrates that low overall discovery probability can coexist with much higher replicability among significant results that reflect genuine effects.
In short, evaluations of credibility need to distinguish among several quantities: the probability of significance across all tests, the probability of significance among true alternatives, the replicability of results selected for significance, and the probability that a significant result is a false discovery. Median discovery probability provides little information about the latter two quantities.
Askarov et al.’s finding of low median power therefore does not by itself imply that economics research lacks credibility. Their own mean-power estimate and the z-curve EDR both suggest an underlying discovery probability of approximately 27%, while z-curve estimates an ERR of approximately 69% and a maximum FDR of approximately 14%. These results indicate substantial selection for statistical significance and considerable room for improvement, but they do not support the conclusion that the low median power of individual estimates, by itself, raises serious doubts about the credibility of the meta-analyzed economics literature.
Conclusion
In conclusion, meta-scientists often point out that extraordinary claims require extraordinary evidence and that academic incentives can reward researchers for making strong claims from weak evidence. Meta-science is not immune to these pressures. The claim that an entire discipline conducts studies with a typical probability of only 7% of rejecting a false null hypothesis is remarkable, if true. However, closer examination shows that this headline figure is a median discovery probability and is not the quantity that predicts the expected number of significant results or the credibility of significant findings. Askarov et al.’s own mean estimate is approximately 27%, closely matching the z-curve EDR, while z-curve estimates substantially higher replicability among significant results and a relatively modest upper bound on the false discovery rate. Thus, the evidence supports concerns about selective reporting and low discovery rates, but it does not support the much stronger conclusion that the low median power estimate by itself raises serious doubts about the credibility of economics research.