Faridani, S. (2026). Testing for underpowered literatures. Journal of Econometrics, 257, 106312. https://doi.org/10.1016/j.jeconom.2026.106312
Abstract
Faridani (2026) introduces a nonparametric deconvolution method for estimating how much the proportion of statistically significant results would increase if studies used larger samples. The method is innovative because it avoids specifying a parametric distribution of true effects and can accommodate publication bias and point masses at zero. However, the estimand is an unconditional increase in significance rates, not statistical power conditional on a true alternative. As a result, a small estimated gain can reflect high power, many true or approximately true null effects, very small nonzero effects, or some combination of these. For research planning and cost-benefit decisions, methods that estimate the latent mixture of null and non-null effects and the power distribution among true effects may therefore provide more directly relevant information.
Review
Faridani develops an innovative method for addressing an important problem in meta-research: determining how much larger samples would have changed the statistical conclusions of a heterogeneous collection of studies. Rather than attempting to estimate the power of individual studies, the paper estimates the counterfactual increase in the proportion of statistically significant results if all sample sizes were multiplied by a common factor. The proposed estimator uses nonparametric deconvolution, accommodates publication bias, and allows substantial heterogeneity in true effects, including point masses at zero.
A major strength of the paper is its recognition that estimating the full distribution of true effects is a much harder inverse problem than estimating a particular functional of that distribution. Faridani therefore focuses on , the expected increase in the proportion of significant results when sample size is increased. This makes the problem statistically more tractable. The simulations suggest that the estimator performs reasonably well across a range of data-generating processes, and the empirical applications to economics and Many Labs provide useful demonstrations.
The interpretation of , however, deserves some qualification. Faridani motivates it as a measure of whether a literature is “undersized” and as information relevant to decisions about whether research funds should be spent on larger samples. This is a coherent estimand, but it is not the same as conventional statistical power, which is defined conditional on a false null hypothesis.
For a literature containing both true null and true alternative hypotheses,
Increasing sample size leaves the first component unchanged. Therefore,
A small value of can therefore have several very different interpretations. Studies testing real effects may already have high power, in which case larger samples are indeed unnecessary. But can also be small because many hypotheses are true or approximately true nulls, or because many true effects are too small for a moderate increase in sample size to help much.
A simple example illustrates the problem. Consider one literature in which every study has a moderate true effect but poor precision. Such a literature may have uniformly low conditional power and would clearly benefit from larger samples. Now compare it with a literature in which half the hypotheses are exactly null and the other half have large effects that are already detected reliably. The second literature may show an equally small or smaller increase in the number of significant results, but for a completely different reason. In the first case, low power is the problem. In the second, many non-significant results are correct indications of little or no effect.
This distinction matters for the funding interpretation. The scientific objective is not simply to maximize the number of results. Ideally, a test should have a high probability of rejecting when a meaningful alternative is true while retaining the nominal Type I error rate when is true. A small increase in the unconditional rejection rate therefore does not by itself show that studies testing real effects have adequate power.
The main uncertainty lies among the non-significant results. If significant results already have high expected replicability, increasing their sample sizes has little practical value. The important question is whether non-significant results are mostly true or approximately true nulls, or false negatives arising from real effects studied with insufficient precision. The answer depends on the latent mixture of and and on the effect-size distribution within the component.
Faridani’s method deliberately avoids recovering this structure, which is part of what makes the estimand tractable. But this also limits its usefulness for planning future sample sizes. A deconvolution estimate of the average increase in significant results does not separately estimate the prevalence of null and non-null effects or the conditional power of studies testing real effects.
Mixture approaches such as z-curve aim to recover more of this latent structure. In particular, they can estimate the expected replicability of significant results and features of the latent power distribution. This permits a more informative separation between the prevalence of true effects and the power of studies conditional on those effects being present. Such information is more directly useful for cost-benefit analyses of future sample sizes than an average increase in significance rates that combines true or approximately true nulls with true alternatives.
A second, more technical point concerns the paper’s characterization of existing mixture methods. Faridani contrasts his nonparametric approach with methods such as Brunner and Schimmack (2020), which he describes as requiring a “specific shape” for the latent distribution. That description is somewhat too strong. A finite mixture on a prespecified grid does impose an approximation, but its purpose is precisely to approximate a wide range of latent distributions by estimating the mixture weights. The more informative distinction is therefore between different forms of regularization of the same inverse problem: finite-mixture approximation versus regularized nonparametric deconvolution.
These qualifications do not detract from the methodological contribution. Faridani provides a clever way to estimate a difficult counterfactual quantity without recovering the full latent distribution. But the resulting estimand answers a narrower question than the one most directly relevant to statistical power and research planning. It tells us, “How many more significant results would larger samples produce?” A more informative planning question is, “Which studies are underpowered because they are testing real effects, and how much would larger samples improve their ability to detect those effects?” Answering that question requires distinguishing the and components rather than averaging across them.