Maier, M., Bartoš, F., & Wagenmakers, E.-J. (2023). Robust Bayesian meta-analysis: Addressing publication bias with model-averaging. Psychological Methods, 28(1), 107–122. https://doi.org/10.1037/met0000405
Take Home Message
A RoBMA estimate near zero is not, by itself, evidence that there is no effect. It is a model-averaged posterior estimate that can be pulled toward zero by the prior for the effect under and then pulled toward zero again by posterior weight on the point-null model. Therefore, a RoBMA estimate near zero does not imply that the true average effect is zero, or even close to zero.
Moreover, if RoBMA also finds substantial heterogeneity, even a true average effect of zero should not be interpreted as evidence that “there is no effect.” It means only that positive and negative population effects balance so that their average is zero. Individual effects may still be substantial. Thus, the point-null hypothesis for the mean is not equivalent to a no-effect hypothesis when effects are heterogeneous.
Finally, model averaging does not guarantee that the resulting average is closer to the truth. It averages estimates from models according to their posterior weights, but all of those models may rely on incorrect or weakly identified assumptions. It is therefore important to examine the estimates from the individual models, the conditional estimate under , the posterior evidence for heterogeneity, and the sensitivity of the result to alternative prior specifications rather than interpreting the model-averaged estimate alone.
Introduction
Psychology has been trying for decades to become a real science. However, the replication crisis in social psychology has revealed that the scientific practices of empirical psychologists produce many results that cannot be replicated in later studies (Open Science Collaboration, 2015). The realization that published results are not as robust as they appear has been called the credibility crisis. In response to the credibility crisis, methodologists and meta-psychologists have debated the causes of replication failures and proposed solutions to the credibility crisis.
Some suggestions have produced real change that increases credibility. Personally, I would argue that the sharing of data has been the biggest positive change. Before, researchers would state that data are available on request, but requests often resulted in apologies why the data could not be shared (my dog ate my USB stick). Now it is expected and sometimes even requested by journals that data are shared openly. That makes it possible to find errors in the statistical analyses and occasionally even evidence of data tampering. However, when open data reproduce the original result, it still remains unclear what the data tell us about a particular hypothesis or theory. The reason is more complex and rooted in the problem that many psychological effects are weak and that studies often have large amounts of sampling error. Sampling error simply means that random factors influence the results and a new sample would produce different results. Although this problem is not new or specific to psychology, it remains unsolved. That is, the same data can be analyzed with methods that produce different answers. One method may conclude that the data support a hypothesis and another one leads to the opposite conclusion. The main problem is that empirical researchers often lack the training (and motivation) to understand why different methods produce different results.
Nil Hypothesis Testing Is Wrong
The reason why different models produce different results with the same data is that they make different assumptions about the data and that they answer different questions. The problem is that researchers have a simple question (e.g., how much does listening to music during studying decrease test performance,” but there is no simple way to answer this question. Statistical methods answer questions that can be answered and researchers then report these answers as if they answered the actual question.
This indirect approach of drawing inferences from data is enshrined in the nil-hypothesis ritual that still dominates psychology. After allocating students to a condition with music or no music, the means are compared and the sampling error (standard error) is computed. The ratio of the two is used as a measure of how compatible the data are with the nil hypothesis that there is absolutely no difference. Exactly zero, zip, nada. If the test statistic is above a conventional criterion value, the nil hypothesis is rejected in favor of the conclusion “Music influences learning!” This ritual has been criticized more often than anybody can count and it doesn’t even tell us whether music increases of decreases performance. It just says, the effect is not nil (zero). Let’s say, the sign is negative. However, we still do not know how much music decreases performance. Maybe it changes a grade by 1 percentage point, maybe it is 10 percentage points. So, the answer to the question “How much does music decrease performance” is “It is not zero.” The answer is correct, but not the answer we wanted.
Even worse, the answer could be false. Psychologists use a criterion value that produces a false positive result in 5% of all tests in which the nil hypothesis is true; the effect is really exactly zero. An error rate of 5% may seem low, but this does not mean that only 5% of the results reported in textbooks are false. The reason is that textbooks focus on the studies that rejected the nil hypothesis for a good reason. The other studies are simply inconclusive and do not tell us anything about direction or strength of an effect, if we use this statistical approach.
If results are selected because they showed an effect, it remains possible that all of the published results are false rejections of a hypothesis. This is not just a hypothetical concern. Some people have argued that published studies report more false results than true results.
In conclusion, the most widely used statistical approach in empirical psychology makes it possible that most published results are false, and the replication crisis suggested that false positives contribute considerably to replication failures.
Bayesian Salvation?
The credibility crisis created an opportunity for Bayesian statisticians who had long criticized nil-hypothesis significance testing. One attraction of Bayesian hypothesis testing is that it promises something conventional significance testing cannot provide: evidence in favor of the nil hypothesis. A nonsignificant -value does not tell us that the nil hypothesis is true. A Bayes factor, in contrast, can compare a model in which the effect is exactly zero with a model in which the effect can take other values.
This sounds like progress. If psychologists have produced too many false positive results, a method that can distinguish evidence for an effect from evidence for no effect seems useful. The problem is that the ability to produce evidence for the nil hypothesis does not guarantee that the evidence is correct. Just as significance testing can falsely reject a true nil hypothesis, Bayesian hypothesis testing can favor a false nil hypothesis.
Whether this happens depends on the assumptions built into the Bayesian models and their priors.
The Problem of Priors
Bayesian statisticians are aware that prior assumptions influence their results, and they are usually explicit about this. However, the terminology can be misleading to users who are not Bayesian statisticians. Some priors are described as “objective” priors. In everyday language, objective sounds as if the prior is neutral and does not influence the result. That is not what the term means.
An objective prior is better understood as a conventional prior that was chosen without considering the specific subject matter. In this sense, it is similar to using an alpha level of .05 in significance testing. Researchers do not choose .05 because there is something special about their particular hypothesis that makes .05 appropriate. They use it because it has become a convention. Likewise, an objective prior is typically based on a general rule that can be applied across many different research questions.
But standardized does not mean neutral. The prior is still an assumption, and it can influence the result. Different priors applied to the same data can produce different posterior estimates and different conclusions. There is nothing mathematically wrong with this. It is an intended feature of Bayesian statistics. The problem arises when results are presented as if they were produced by the data alone.
A Bayesian result is always conditional on both the data and the prior assumptions used to interpret them. Thus, “objective” does not mean “assumption-free.” It means that the assumption was chosen by a general convention rather than by considering the particular phenomenon being studied.
This is similar to other potentially misleading terms in statistics. “Statistically significant” does not mean that an effect is important. In the same way, an “objective prior” does not mean that the prior has no influence on the result. The technical meaning may be clear to statisticians, while the everyday meaning encourages users to infer more than the term actually implies.
The criticism developed here is therefore not a criticism of Bayesian statistics in general. Bayesian statistics allows researchers to use many different priors, including priors informed by substantive knowledge. The concern is what happens when conventional priors are used without considering whether they make sense for the research question at hand.
One particularly important choice is whether one specific value should receive special prior probability. In Bayesian hypothesis testing, that value is often zero. A model in which the effect is exactly zero is compared with a model in which the effect can take a range of other values. This makes it possible to do something that conventional significance testing cannot do: provide evidence in favor of a nil hypothesis.
That sounds like an important improvement. A nonsignificant p-value does not provide evidence that the effect is zero, whereas a Bayes factor can favor a model in which the effect is fixed at zero.
But this advantage creates the possibility of the opposite error. Conventional significance testing can falsely reject a true nil hypothesis. Bayesian hypothesis testing can favor the nil hypothesis when the true effect is not zero.
Suppose, for example, that psychotherapy has a small positive average benefit. A Bayesian analysis could nevertheless favor a model in which the average effect is exactly zero. The calculation may be mathematically correct, while the conclusion is wrong because the prior and model assumptions were poorly suited to the problem.
Before the credibility crisis, psychologists were highly concerned about low statistical power and failures to detect real effects (Cohen, 1994). More recently, attention has shifted heavily toward false positive results. Both types of error matter. A method that reduces false positives by substantially increasing the risk of falsely concluding that real effects are absent is not necessarily an improvement.
The important question is therefore not whether Bayesian statistics is better or worse than conventional significance testing. The important question is how strongly the conclusion depends on the prior assumptions that were chosen and whether those assumptions make sense for the research question.
This issue becomes especially important for Robust Bayesian Meta-Analysis because its default analysis gives zero a privileged role in more than one way.
Meta-Analysis
Disagreement among statisticians often focuses on the problem of drawing conclusions from a single study. How much did racism and sexism influence the outcome of the 2024 U.S. presidential election? A single study cannot provide a definitive answer, and we cannot repeat the election under identical conditions.
Psychologists are often in a better position because many of the effects they study can be examined repeatedly. Does psychotherapy work? This question has been studied for decades and will continue to be studied. So, it may seem that we can simply combine the results of all available studies in a meta-analysis and get a better answer.
Unfortunately, combining many studies is not as simple as it sounds.
The first problem is that effects vary. Psychotherapy may work differently for different disorders, treatments, populations, therapists, and circumstances. There is therefore no single effect size that describes every situation. At best, a meta-analysis can estimate the average effect and how much effects vary around that average.
The second problem is publication bias. Studies with positive or statistically significant results are more likely to be published than studies with negative or inconclusive results. As a result, the published literature can give an exaggerated impression of the average effect.
This problem was largely ignored for many years, and traditional meta-analyses sometimes produced inflated effect-size estimates, including apparently impressive evidence for implausible phenomena such as extrasensory perception. In response, statisticians have developed a variety of methods that try to detect publication bias and estimate what the average effect might have been if all studies had been available.
Unfortunately, this is a difficult problem. Simulation studies show that publication bias can be hard to detect, especially when it is modest. More importantly, different correction methods can produce very different answers from the same data. One method may estimate an average effect of zero, while another estimates an effect of .20.
The reason is that the methods make different assumptions about how studies came to be published and about the distribution of the missing results. When those assumptions are approximately correct, a method may recover the true effect reasonably well. When they are wrong, the estimate can be badly biased.
The crucial difficulty is that the missing studies are, by definition, missing. The observed data often contain too little information to determine which assumptions about the missing results are correct. Consequently, publication-bias adjustments are not simply estimates produced by the data. They are estimates produced by the data plus assumptions about the data we do not see.
Properly interpreted, the results are therefore conditional: “If the assumptions of Model A are approximately correct, the average effect is close to zero. If the assumptions of Model B are approximately correct, the average effect is closer to .20.
Robust Bayesian Meta Analysis (RoBMA)
Robust Bayesian Model Averaging (RoBMA) offers a solution to the problem that different meta-analytic models can produce different answers. RoBMA fits many models to the same data, gives more weight to models that receive more support from the data, and then averages estimates across these models.
However, RoBMA can move estimates toward zero in two different ways. First, even models that allow for a nonzero average effect use a prior distribution for the size of that effect. When the data provide weak information, this prior can pull the conditional estimate toward its prior mean. For example, a true average effect of .20 might produce a conditional estimate of only .10 if the prior favors smaller effects.
Second, RoBMA also includes models in which the average effect is fixed at exactly zero. When RoBMA remains uncertain about whether the average effect is zero or nonzero, these models receive posterior weight in the final average. If the conditional estimate from the nonzero-effect models is .10 and the zero and nonzero models receive roughly equal weight, the final model-averaged estimate will be approximately .05.
Thus, a true average effect of .20 can first be reduced to .10 by the prior within the nonzero-effect models and then reduced again to .05 by averaging this estimate with models that fix the average at zero. Both steps are mathematically legitimate Bayesian operations. But neither step guarantees that the estimate is moving closer to the true value.
This procedure produces a single number, but it does not make disagreement among the models disappear. If one plausible model produces an estimate near zero and another produces an estimate near .20, this disagreement tells us that the answer depends strongly on assumptions. Averaging the estimates does not resolve that problem. Adding further weight to models that fix the average at zero may move the final estimate closer to zero, but there is no reason why closer to zero must mean closer to the truth.
RoBMA therefore does not eliminate the uncertainty created by conflicting models. It summarizes this uncertainty by averaging across them. The danger is that users may interpret the resulting single estimate as if the disagreement had been resolved. In some situations, the disagreement among individual models is more informative than their average because it reveals how strongly the conclusion depends on assumptions.
Default RoBMA
Default RoBMA gives zero a privileged role in two different ways.
The first prior concerns the size of the average effect when an effect is assumed to exist. RoBMA specifies a distribution of plausible effect sizes, and the default distribution is centered at zero. Before seeing the data, small effects are therefore considered more plausible than large effects, with positive and negative effects treated symmetrically. When the data contain strong information about the average effect, this prior has little influence. When the data contain weak information, however, the estimate remains closer to the center of the prior. Thus, even conditional on the assumption that there is an effect, the default prior can move the estimated effect toward zero.
There is nothing inherently Bayesian about centering this prior at zero. A researcher could specify a prior centered at .10, .20, or some other value. The default of zero is a conventional choice made without substantive information about the particular phenomenon being studied. It may be reasonable in some applications, but it is not neutral: changing the center of the prior can change the estimate when the data are weakly informative.
The second prior gives zero an even more privileged role. RoBMA distinguishes between models in which the average effect is estimated and models in which the average effect is fixed at exactly zero. By default, these two possibilities receive equal prior probability.
A 50:50 split can appear neutral because there are two hypotheses and each receives half of the prior probability. But this appearance is misleading. Zero is only one possible value of the average effect. There is no mathematical reason why exactly zero rather than .01, -.01, or .10 should receive a special probability of 50%. We could construct the same Bayesian model with a spike at .10 and give that value half of the prior probability. The resulting model would then be pulled toward .10 rather than toward zero. Bayesian mathematics permits either choice. The reason for privileging zero must therefore come from the substantive interpretation of zero, not from probability theory itself.
One justification for giving zero special status is parsimony. This makes sense for some statistical problems. If a regression coefficient is exactly zero, the corresponding predictor can be removed from the model. The zero coefficient therefore represents a genuinely simpler model: the predictor does not matter.
This logic is much less convincing for the average of heterogeneous effects in a meta-analysis. If some effects are positive and others are negative, an average of zero does not mean that nothing is happening. It merely means that the positive and negative effects cancel on average. Treating this as a no-effect model would be like calculating the average regression coefficient across many predictors, finding that the positive and negative coefficients average to zero, and concluding that all of the predictors can be ignored.
Thus, even if a point mass at zero can be justified for a single effect or a single predictor, the same parsimony argument does not automatically justify giving special status to a zero average when the individual effects vary.
Nor is there empirical evidence that the average effects studied in psychology have a 50% probability of being exactly zero. The 1:1 prior is a conventional default, not an empirical estimate of how often nil hypotheses are true. Other prior probabilities can be chosen, but there is no generally accepted value that follows from the data themselves.
More importantly, a prior probability for the nil hypothesis is not needed simply to estimate the average effect size. It is needed because RoBMA also asks a categorical question: should the data be attributed to a model in which the average effect is exactly zero or to models in which the average effect is allowed to differ from zero?
These are different questions. Researchers may want to know whether the data discriminate between a point-null model and an alternative model. But the usual goal of a meta-analysis is more straightforward: What is the best estimate of the average effect, and how uncertain is that estimate?
The distinction matters because the answer to the categorical question feeds back into the reported model-averaged effect size. If RoBMA remains uncertain about whether the average is exactly zero, posterior probability remains on the zero model. The conditional estimate from the models that allow an effect is then averaged with zero. Consequently, uncertainty about whether the nil hypothesis is true becomes additional shrinkage of the estimated effect toward zero.
This means that two different prior assumptions can work together. The prior for the size of the effect can pull the conditional estimate toward zero, and the prior probability assigned to the point-null model can then pull the model-averaged estimate toward zero again.
Neither operation is mathematically incorrect. But neither guarantees movement toward the truth.
A Simulation Example
Many meta-analyses in psychology show small average effects together with substantial heterogeneity (van Erp et al., 2017). I used a simple simulation to examine whether RoBMA can produce false evidence for the nil hypothesis under these conditions.
True effect sizes were drawn from a normal distribution with an average effect of and a standard deviation of .40. Thus, the true average effect was clearly nonzero, but individual effects varied widely and could be either positive or negative.
Rather than simulating particular sample sizes, I simulated the standard error of each study directly from a uniform distribution ranging from .05 to .30. True effect sizes and standard errors were independent.
This setup deliberately satisfied two important assumptions of models included in RoBMA. First, true effect sizes followed a normal distribution, which is the population distribution assumed by the random-effects selection models. Second, true effect sizes were independent of their standard errors, so there was no genuine small-study effect of the kind modeled by PET and PEESE.
Publication bias was introduced in a particularly simple way. Only studies with positive observed effect sizes were retained. There was no additional selection for statistical significance and no p-hacking. Selection depended only on the direction of the observed effect.
This form of publication bias is directly relevant to RoBMA because its step-function selection models can include a cutoff at , allowing studies with effects in one direction to have different publication probabilities from studies with effects in the opposite direction. Thus, the simulation did not confront RoBMA with a form of selection that lies outside the model space.
Ironically, the model designed to handle directional selection is central to the problem examined here. Given a normal distribution of heterogeneous effects, an observed literature containing only positive results implies that negative results are missing. RoBMA therefore has to reconstruct the unseen negative part of the distribution.
The difficulty is that the observed positive studies do not uniquely determine where the center of the complete distribution lies. A population centered at .20 with substantial heterogeneity and many missing negative results can generate the observed data. But a population centered much closer to zero, combined with a somewhat different amount of directional selection, can also fit the observed positive tail. Because RoBMA includes models that fix the average effect at zero and gives these models positive prior weight, this ambiguity can move the model-averaged estimate toward zero even when the true average is .20.
This is the key identification problem. The data contain a great deal of information about the studies that were observed, but much less information about the studies that were not observed. As a result, RoBMA can remain uncertain about the true average even when hundreds of studies are available.
That uncertainty creates room for the priors to matter. The prior under the alternative hypothesis can pull the conditional estimate toward its own center, and the separate point-null prior can then pull the model-averaged estimate toward zero again.
The simulation is therefore intentionally favorable to RoBMA in several respects. The true effect distribution is normal, there is no relationship between true effect sizes and standard errors, there is no p-hacking, and the publication process consists of simple directional selection that RoBMA explicitly allows for. Nothing in this setup would lead an applied researcher to expect that default RoBMA should frequently provide evidence for an average effect of exactly zero when the true average effect is .
Simulation Results
Across 100 simulations, the average RoBMA estimate of the average effect size was only d = .065, although the true average was . All credible intervals included zero, whereas only 81% included the true value of .20. Thus, the intervals consistently treated zero as plausible, while failing to include the true effect in 19% of the simulations. The uncertainty intervals therefore do not provide convincing evidence for the true effect and are clearly shifted toward zero.
RoBMA also reports Bayes factors that compare models in which the average effect is exactly zero with models in which it is not. Bayes factors are continuous measures of evidence and do not require a decision to accept or reject the nil hypothesis. In practice, however, conventional cutoffs are often used to describe evidence as favoring one hypothesis or the other. A Bayes factor greater than 3 is commonly interpreted as moderate evidence.
Using this criterion, RoBMA found moderate evidence for a real average effect in only 6% of the simulations. More importantly, it found moderate evidence for the false hypothesis that the average effect was exactly zero in 31% of the simulations. Thus, RoBMA did more than remain uncertain about a small true effect. In about one third of the simulations, it provided evidence in the wrong direction.
RoBMA can also report a conditional estimate that excludes models in which the average effect is fixed at zero. Across simulations, this conditional estimate averaged d = .130. Thus, even within models that assume an effect exists, the estimate had already been pulled substantially from the true value of .20 toward zero.
The default model-averaged estimate was still lower, because RoBMA further down-weights the conditional estimate according to the posterior probability that an effect exists. On average, this posterior probability was only about 40%, below its prior value of 50%. The average model-averaged estimate is not simply , because the conditional estimate and posterior probability vary together across simulations. Nevertheless, the mechanism is clear: substantial posterior probability remains on models that fix the average effect at exactly zero, and these models contribute zero to the final average.
The estimate is therefore moved toward zero twice. First, within models that allow an effect, the true average of .20 is reduced to a conditional estimate of .126. Second, averaging these estimates with models that fix the effect at zero reduces the final reported estimate further, to .064.
The different model families reveal an additional problem. When RoBMA is restricted to the step-function selection models, it reproduces the near-zero estimate. When it is restricted to PET/PEESE regression models, however, it produces an estimate of approximately .
At first sight, the PET/PEESE estimate may appear to be a severe overestimate because the true average across all effects is .20. But PET/PEESE does not model selection by direction. It only adjusts for the part of the observed effect that is associated with standard error—the small-study-effect pattern that often results from selective publication of statistically significant findings.
In this simulation, selection was based on direction rather than significance: only positive observed effects were retained. This selection induced only a small correlation between effect size and standard error, about . The mean observed effect among the selected studies was approximately , whereas the mean true effect among those selected studies was approximately . PET/PEESE detected the small association with standard error and reduced the estimate slightly, from .44 to .42. Thus, under this selection mechanism, PET/PEESE effectively estimates the average effect among the surviving positive studies rather than reconstructing the mean of the complete population.
The contrast between the model families is therefore striking. The step-function models produce an estimate close to zero, whereas PET/PEESE produces an estimate around .42. These are not simply two noisy estimates of the same quantity. They reflect fundamentally different assumptions about what happened to the missing studies and, in this simulation, effectively refer to different quantities.
Ironically, an equal arithmetic average of the step-function estimate of about .06 and the PET/PEESE estimate of about .42 would be approximately .24, close to the true value of .20. But this would be an accident, not a meaningful solution. Nobody is interested in the average of the mean across all effects and the mean among only the positive effects.
RoBMA did not average these two estimates equally. It assigned essentially zero weight to the PET/PEESE models and therefore reported an estimate dominated by the step-function models. The fact that this estimate happened to be much farther from the truth than the PET/PEESE estimate illustrates an important limitation of model averaging: giving more weight to one model does not guarantee that its estimate is closer to the true value.
In sum, this example demonstrates two distinct problems. First, RoBMA can provide evidence that the average effect is exactly zero when the true average is a small positive effect. Second, averaging across models with different assumptions does not guarantee a more accurate or even a more meaningful estimate. In this case, the disagreement between estimates of approximately .06 and .42 is itself important information. It reveals how strongly the answer depends on assumptions about the missing studies. Averaging that disagreement away does not resolve the uncertainty.
Alternative Solution: Mindful Meta-Analysis
RoBMA was developed to address a real problem: different methods for correcting publication bias can produce dramatically different answers from the same data. This blog post shows why automatically averaging these answers does not solve the problem. The models disagree because they make fundamentally different assumptions about how the observed data were generated and about the results that were not observed. Combining their estimates into a single number does not resolve these disagreements. It can merely hide them behind an algorithmic average.
This is similar to the problem Gigerenzer identified in the ritualistic use of nil-hypothesis significance testing. The problem was not the mathematics of significance testing. It was the replacement of scientific reasoning by an automatic procedure: calculate a p-value, compare it with .05, and report the appropriate conclusion. Replacing this ritual with another automatic procedure—fit many Bayesian models, average them, and report one number—does not solve the fundamental problem. Data analysis requires researchers to think about what their models assume and what question they are trying to answer.
In this example, asking about the true average for a hypothetical population of effect sizes that includes negative values when only positive values are observed is the wrong question. The same positive distribution could arise in a fundamentally different way. Imagine that effects were generated directly from a distribution that contained only positive values. No negative results would have been suppressed because no negative effects ever existed. The observed data could look essentially the same, but reconstructing a large number of hypothetical negative results would now be wrong.
This is the fundamental identification problem. From the observed positive results alone, the model cannot know whether negative results once existed and were suppressed or whether the population that generated the data simply contained few or no negative effects. The answer comes from assumptions about the missing part of the distribution, not from direct observation.
In this specific case, a more meaningful question is whether the positive observed results all tested a true nil hypothesis, which is already rejected by the large estimate of heterogeneity that implies at least some of the positive results were obtained with real effects. In this case, the PET-PEESE estimate actually provided the correct answer about the average effect size of the observed results. In contrast, users who focus on the average estimate of RoBMA might arrive at the false conclusion that the true effect size is zero, which ignores the evidence of heterogeneity. When studies are heterogeneous, there is no single effect size to estimate. Thus, even the misleading estimate of zero would require further examination of the positive results as a next step.
In conclusion, meta-analysis requires a lot more thinking than published meta-analyses that focus on the average effect size estimate imply. The devil is not only in the defaults of meta-analytic models, it is also hidden in large estimates of heterogeneity that remains unexplained. If mindless use of p-values got us into the mess, mindless use of Bayesian model averages is not going to get us out of it.





