Assessing Nudge Impact: A Comprehensive Second- Order Monster- Analysis

Hu, B., Z. Xia, Q. Guo, C. Lu, S. Constantino, and X. Ju. 2025. “ Assessing Nudge Impact: A Comprehensive Second-Order Meta-Analysis.” Journal of Behavioral Decision Making 38, no. 5: e70053. https://doi.org/10.1002/bdm.70053.

Post-Publication Review

Hu et al. present a second-order meta-analysis of 14 meta-analyses covering 1,638 studies and nearly 30 million participants. They report an unadjusted mean effect of d=.27d=.27, while PET-PEESE reduces the estimate to approximately zero.

The main problem is that adding more and more studies does not solve the basic problem of extreme heterogeneity. The authors report I2=99.89%I^2=99.89\%, indicating that the included meta-analyses differ enormously in their estimated effects. Under these conditions, the question “What is the effect of nudging?” is not well defined. Defaults, reminders, food placement, information labels, and many other interventions are not repeated estimates of one common treatment effect.

The publication-bias results do not resolve this problem. In fact, they disagree. The conventional random-effects estimate is d=.27d=.27. RoBMA gives d=.29d=.29, 95% CI [.15, .43], with τ=.26\tau=.26. Trim-and-fill still gives d=.27d=.27, and using adjusted estimates from component meta-analyses gives d=.28d=.28. Only PET-PEESE produces the near-zero estimate of d=.003d=.003–.004.

It is therefore misleading to foreground the PET-PEESE result as if different methods converge on zero. RoBMA already averages over PET-PEESE, selection, and no-bias models, yet its model-averaged estimate remains positive. Moreover, PET-PEESE relies on the relation between effect size and standard error, which can be confounded by genuine heterogeneity. In a literature where small and large studies differ in intervention type, setting, population, and outcome, that relation need not reflect publication bias alone. A separate weight-function estimate would have been informative, but none is reported.

More fundamentally, even a perfectly estimated grand mean would tell us little. The authors themselves show that effect sizes differ substantially across categories and domains; for example, decision-structure nudges have a descriptive estimate of d=.40d=.40, environmental nudges d=.45d=.45, and food nudges d=.33d=.33. Their moderator tests are underpowered and do not explain the heterogeneity, but failure to explain heterogeneity does not make the grand mean meaningful.

The article therefore does not tell us much that is new about nudging. We already knew that published effects are heterogeneous and that publication bias is a concern. Pooling ever larger numbers of heterogeneous studies does not answer the scientifically important questions: Which nudges work? How large are their effects? Under what conditions do they work?

The conclusion that the “true effect of nudging” is zero has the appearance of precision, but it is closer to answering an ill-defined question. The problem is not whether the grand mean is d=.27d=.27 or d=.004d=.004. The problem is that there is no single effect of nudging to estimate.

Leave a Reply