How Robust is RoBMA?

Robust Bayesian Meta-Analysis, or RoBMA, was developed in E. J. Wagenmakers’ research group as a method for evaluating meta-analytic evidence while allowing for the possibility that published effects are exaggerated by publication bias. An important feature of RoBMA is that it does not merely estimate an effect size; it also assigns evidence to models in which the average effect is exactly zero.

This approach builds on a form of Bayesian hypothesis testing that Wagenmakers and colleagues have long promoted as more balanced than conventional null-hypothesis significance testing. In standard significance testing, researchers can reject the nil hypothesis or fail to reject it, but a nonsignificant result does not itself constitute evidence that the effect is zero. Bayes factors, by contrast, can quantify relative evidence for either the nil or an alternative hypothesis.

This form of Bayesian hypothesis testing dates back to Harold Jeffreys’ Theory of Probability (1939), but it remained far outside the mainstream of psychological research for decades. It became highly visible in psychology after Bem (2011) reported a series of statistically significant experiments purporting to demonstrate extrasensory perception. Wagenmakers et al. (2011) reanalyzed Bem’s (2011) results with default Bayes factors and argued that the evidence for psi was weak or nonexistent despite the significant pp-values. However, the Bayes Factors also failed to produce clear evidence that ESP does not exist.

Robust Bayesian Meta-Analysis, RoBMA for short, was developed to conduct Bayesian meta-analyses that take publication bias into account and can provide evidence for the nil hypothesis that the average true effect is zero. A Bayes-Factor that does not favor the broad alternative hypothesis of an effect is sometimes interpreted as “No evidence for an effect” (Maier et al., 2022). However, this conclusion is not warranted if the model also estimates heterogeneity around this average. Heterogeneity implies that some estimates reflect true effects. Moreover, the average effect size of positive results is implied to be greater than zero (Schimmack, 2026).

The main problem with RoBMA is that it s built on top of existing bias-correction models with strong assumptions, specifically step-function selection models (Vevea & Hedges, 1995; Hedges & Vevea, 1996). The most critical assumption of this model is that population effect sizes are normally distributed. When this assumption is violated because the dataset contains few negative results, the model assumes that selection bias removed negative results and produces a lower estimate closer to zero. In extreme cases, it can estimate that the mean is zero and find evidence for a false nil hypothesis about the mean.

When RoBMA was used with data from a meta-analysis of nudging interventions, the model adjusted the average effect of approximately d=.43d=.43, to d = .07. However, another bias correction methods that do not assume a normal distribution, produced an estimate of d= .44 (Schimmack, 2026b).

RoBMA Analysis of Government Nudging Studies

DellaVigna and Linos (2022) used results from large government nudge trials and compared them with results from broadly comparable academic studies. The difference was striking. In the academic literature, nudges changed behavior by about 8.7 percentage points on average. Thus, if 20% of people performed a behavior in the control condition, a typical effect of this magnitude would increase the rate to about 29%. In the government trials, however, the average effect was only about 1.4 percentage points. In the large government studies, this finding was still statistically significant, but the same effect size might be too weak in other studies to reject the nil hypothesis. Thus, the results of this study seem to be consistent with RoBMA’s bias-corrected estimate. However, an alternative explanation is that there government studies have smaller effect sizes because they are different from the nudging manipulations in academic studies. To explore these competing hypothesis, I analyzed the data with RoBMA and zcurve3, a model that does not assume a single normal distribution of population effect sizes.

The authors of RoBMA already analyzed these data, but they did not use RoBMA (Maier et al., 2024), although a reviewer suggested it. They argued that it would be inappropriate to use RoBMA with these data because the data do not fit a single normal distribution. This is an important caveat because it makes the distribution assumption salient and because the authors applied RoBMA to the academic meta-analysis without testing the distribution assumption.

Here I applied RoBMA to data that are known to violate the distribution assumption. The result was similar to the one for the academic meta-analysis: evidence for an overall mean effect was inconclusive and slightly favored the null model (inclusion BF = 0.46). The model-averaged estimate of the mean effect was μ=0.016\mu = 0.016, 95% credible interval [-0.413, 0.502]. In other words, “no evidence for an average effect in government nudging studies.” Again, this finding does not mean that nudging interventions do not work. Given large heterogeneity around the average, The analysis provided overwhelming evidence for heterogeneity in true effects (inclusion BF > 14,999), with an estimated heterogeneity parameter of τ=3.34\tau = 3.34, 95% credible interval [3.07, 3.63]. However, the results imply that nudging studies also often backfire, but that these negative results are missing, publication bias (inclusion BF > 14,999).

Figure 1 from Maier et al. (2024) shows the problem for the weight-function selection model underlying RoBMA. The distribution has a long positive tail that does not fit a single normal distribution (blue). A model that only fits the right side and attributes the low frequency of negative results to selection bias (red line). However, even this model does not fit the data well, but RoBMA has no other way to represent the data.

However, we can help RoBMA by removing extreme values from the data on both sides. However, because there are many more positive extreme values than negative extreme values, the exclusion of extreme values lowers the average of the observed data from 1.38 percentage points (pp) to 0.74 pp. In contrast, the median did not change much (0.50 vs. 0.42 pp).

Applying RoBMA to these data changed the estimate of the average effect size dramatically. The Bayes Factor now clearly rejected the nil hypothesis, BF > 14,999 (the maximum value in RoBMA). The estimated average was 0.58, only slightly lower than the simple average of 0.74 pp. This estimate also does not take the extreme positive results into account. The main point, however, is that RoBMA underestimated the average effect size when a long tail of positive results was included because it imagined a long tail of matching negative results based on the symmetry of the normal distribution. As Maier et al. (2024) already stated, this means RoBMA should not be used when data are likely to have a different distribution.

Maier et al. (2022) computed their RoBMA result ignoring the clustering of standard errors — as stated in their own figure caption. This is not a minor detail. The 448 effect sizes come from 212 studies, and treating dependent estimates as independent inflates the apparent precision that the bias correction then acts on. Without clustering, RoBMA returns a mean of essentially zero and evidence against an effect under every bias specification I tried — the full model-averaged ensemble, and a selection-model-only version with PET-PEESE removed (μ = 0.02, BF favoring the null). Once the clustering is modeled — using the multilevel extension of RoBMA (Bartoš, Maier & Wagenmakers, 2026) — the evidence reverses decisively (inclusion BF > 14,999) to a clearly positive mean, between d = 0.22 and d = 0.34 depending on the bias model. The published conclusion of “no evidence for nudging” is an artifact of ignoring the dependence structure of the data.

How Robust is RoBMA?

In statistics, robustness traditionally refers to methods that remain reliable when assumptions are violated or when data contain outliers or other atypical observations (e.g., Box, 1953; Huber, 1964; Tukey, 1960). Calling RoBMA “robust” may therefore give users the impression that it has these properties. However, robust in RoBMA refers primarily to averaging across a predefined set of meta-analytic models, not to robustness against misspecification of assumptions those models share. Bayesian model averaging can hedge against uncertainty about which model in the set is correct; it cannot protect against an assumption that every model in the set makes.

That distinction is important because RoBMA, like every statistical model, can produce misleading estimates when its assumptions are poorly suited to the data. In the present analyses, results depend strongly on assumptions about the distribution of population effect sizes and the treatment of extreme or highly heterogeneous effects. Other plausible models can produce substantially different conclusions.

The practical lesson is therefore that its results should not automatically be treated as the most robust or most credible estimate. When different models give different answers, the scientifically important task is to examine why they disagree and which assumptions are plausible for the data at hand.

Leave a Reply