Blogging about statistical power, replicability, and the credibility of statistical results in psychology journals since 2014. Home of z-curve, a method to examine the credibility of published statistical results.
Show your support for open, independent, and trustworthy examination of psychological science by getting a free subscription. Register here.
“For generalization, psychologists must finally rely, as has been done in all the older sciences, on replication” (Cohen, 1994).
DEFINITION OF REPLICABILITY: In empirical studies with sampling error, replicability refers to the probability of a study with a significant result to produce a significant result again in an exact replication study of the first study using the same sample size and significance criterion (Schimmack, 2017).
See Reference List at the end for peer-reviewed publications.
Mission Statement
The purpose of the R-Index blog is to increase the replicability of published results in psychological science and to alert consumers of psychological research about problems in published articles.
To evaluate the credibility or “incredibility” of published research, my colleagues and I developed several statistical tools such as the Incredibility Test (Schimmack, 2012); the Test of Insufficient Variance (Schimmack, 2014), and z-curve (Version 1.0; Brunner & Schimmack, 2020; Version 2.0, Bartos & Schimmack, 2021).
I have used these tools to demonstrate that several claims in psychological articles are incredible (a.k.a., untrustworthy), starting with Bem’s (2011) outlandish claims of time-reversed causal pre-cognition (Schimmack, 2012). This article triggered a crisis of confidence in the credibility of psychology as a science.
Over the past decade it has become clear that many other seemingly robust findings are also highly questionable. For example, I showed that many claims in Nobel Laureate Daniel Kahneman’s book “Thinking: Fast and Slow” are based on shaky foundations (Schimmack, 2020). An entire book on unconscious priming effects, by John Bargh, also ignores replication failures and lacks credible evidence (Schimmack, 2017). The hypothesis that willpower is fueled by blood glucose and easily depleted is also not supported by empirical evidence (Schimmack, 2016). In general, many claims in social psychology are questionable and require new evidence to be considered scientific (Schimmack, 2020).
Each year I post new information about the replicability of research in 120 Psychology Journals (Schimmack, 2021). I also started providing information about the replicability of individual researchers and provide guidelines how to evaluate their published findings (Schimmack, 2021).
Replication is essential for an empirical science, but it is not sufficient. Psychology also has a validation crisis (Schimmack, 2021). That is, measures are often used before it has been demonstrate how well they measure something. For example, psychologists have claimed that they can measure individuals’ unconscious evaluations, but there is no evidence that unconscious evaluations even exist (Schimmack, 2021a, 2021b).
If you are interested in my story how I ended up becoming a meta-critic of psychological science, you can read it here (my journey).
References
Brunner, J., & Schimmack, U. (2020). Estimating population mean power under conditions of heterogeneity and selection for significance. Meta-Psychology, 4, MP.2018.874, 1-22 https://doi.org/10.15626/MP.2018.874
Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17, 551–566 http://dx.doi.org/10.1037/a0029487
Schimmack, U. (2020). A meta-psychological perspective on the decade of replication failures in social psychology. Canadian Psychology/Psychologie canadienne, 61(4), 364–376. https://doi.org/10.1037/cap0000246
Bargh, Chen, and Burrows (1996) reported that priming students with words related to old age made them walk more slowly, with effect sizes above one standard deviation (d = 1.04 and .79) across two studies of just 30 participants each. The finding became a cornerstone of behavioral-priming research and a famous casualty of the replication crisis after Doyen et al. (2012) failed to reproduce it. The Doyne et al. (2012) article came at the right time to cause a paradigm shift, but it was not the first replication failure.
The first replication failure appeared only a couple of years after the 1996 article: Dijksterhuis et al. (1998) found the same comparison at roughly a quarter of a standard deviation (d = .25 and .29), nonsignificant, and reinterpreted the disappearance as a side condition for a new contrast effect. Because social psychologists tracked the sign of an effect rather than its magnitude, a series of studies that did not reproduce Bargh’s result were published as successful demonstrations of moderators instead of as replication failures.
By 2012, elderly priming had been reported in 8 articles with 12 studies and 15 tests. I show that a meta-analysis at this time would have shown a wide prediction interval with possible effect sizes ranging from -0.1 to +1.0. Thus, Doyen et al.’s replication failure was entirely consistent with the full existing evidence. It was therefore entirely reasonable for Kahneman (2012) to ask for new and stronger evidence that Bargh never delivered. Since then, independent preregistered studies large failed to replicate past effects and priming theorists have largely abandoned priming research. The death of the priming paradigm provides a valuable lessons about the need to build theories on robust empirical foundations to avoid investing resources on phenomena that do not exist.
The Long Version
John A. Bargh studied with Robert Zajonc at the University of Michigan, earning his Ph.D. in 1981. Zajonc was an early champion of unconscious processes and used masked, subliminal presentations to study influences occurring outside conscious awareness. At New York University, Bargh developed a related program on automatic social cognition. At the time, studies from several laboratories suggested that stimuli presented outside awareness could influence feelings, judgments, and immediate reactions to other stimuli. Some of this evidence has since been challenged: meta-analyses find many subliminal effects hard to replicate, and concerns have been raised about publication bias and about whether participants were ever fully unaware of the masked stimuli.
Bargh matters for the history of psychology because his 1996 article made a much stronger claim. Rather than showing that unnoticed stimuli can shift immediate perceptions or evaluations, Bargh and colleagues claimed that activating a social concept could have a lasting influence on overt behavior without people being aware of that influence. In their famous elderly-priming experiment, students completed a scrambled-sentence task in which, in one condition, several words were related to old age and, in the control condition, were not. The words were visible, but participants were presumably unaware of the manipulation’s purpose or its possible effect on their behavior. Told to go to another room for a second study, participants then had their walking speed down the hallway measured as the real dependent variable. Bargh et al. reported that a few old-age words made students walk more slowly than controls.
The difference was not small. Effect sizes for two-group differences are often expressed in standard-deviation units. Cohen (1988) classified d = .50 as a medium effect and d = .80 as large. For comparison, an IQ test has a standard deviation of about 15 points, so half a standard deviation is 7.5 IQ points. In Bargh et al.’s first study, Experiment 2a, the effect was slightly larger than a full standard deviation, d = 1.04 — an exceptionally large effect for such a subtle manipulation.
Psychologists rarely run direct replications, but Bargh et al.’s article contained one. Experiment 2b used essentially the same procedure; again, primed participants walked significantly more slowly, and the effect stayed large, d = .79. The article thus appeared to offer unusually convincing evidence: two independent studies, same procedure, both producing significant and large effects.
The 1996 article spawned a large literature using primes such as money or God to influence behavior. Then, in 2012, Doyen et al. reported that they could not replicate the elderly-priming effect and suggested the original findings might be, at least partly, methodological artifacts. Walking speed in the original studies was timed manually with a stopwatch. And although the person timing walking speed was blind to condition, Doyen et al. noted it was unclear whether the experimenter who administered the priming task was also blind. If experimenters knew which participants had been primed, their expectations could have shaped participants’ behavior.
Doyen et al. tested this in a second experiment by manipulating experimenters’ expectations directly: some were led to expect primed participants to walk more slowly, others to walk faster. Objectively measured walking speed tracked those expectations — the elderly-prime effect appeared only when experimenters expected slowing — and the manual stopwatch measurements tracked them even more strongly. Doyen et al. thus provided experimental evidence that researchers’ expectations could produce the outcome of a behavioral-priming study.
Bargh responded in March 2012 with a Psychology Today post titled “Nothing in Their Heads.” He opened by noting that the 1996 finding had been theoretically predicted and fit a growing literature on automatic influences on behavior, but much of the post attacked the quality of Doyen et al.’s work and the peer review at PLOS ONE. Knowledgeable social-psychology editors and reviewers, he argued, would have caught the methodological problems; he had not been asked to review the paper, and implied that had he done so, he would have recommended rejection.
The tone drew wide criticism — Srivastava (2012), for one, characterized the episode as Bargh going “bananas.” The attack on the journal prompted a reply from the PLOS ONE editors on Bargh’s blog post (cf. Hodgkinson, 2012; Srivastava, 2012). Bargh later expressed regret over the tone and took the post down (Bartlett, 2013).
Bargh’s Defense of His Results in 2012
Bargh’s first substantive defense concerned experimenter effects. He stated that the experimenter in the original study had been blind to hypothesis and condition, and that a different person, also blind to condition, measured walking speed. If accurate, this substantially weakens Doyen et al.’s suggestion that experimenter expectations produced the original findings. But it does nothing for the more basic result: Doyen et al. used a larger sample and objective measurement and still failed to reproduce the effect.
Bargh therefore proposed several procedural differences to explain the replication failure. First, he argued that Doyen et al. had drawn participants’ attention to walking by telling them to “go straight down the hall when leaving,” and that making an automatic behavior conscious could eliminate the priming effect. But Doyen et al.’s article contains no such instruction; it says only that participants were “clearly directed to the end of the corridor” — and Bargh et al.’s own procedure had likewise directed participants toward the elevator down the hall. It is unclear the difference existed at all.
Second, Bargh questioned the strength of the manipulation: Doyen et al. put an elderly-related word in all 30 scrambled-sentence items, and Bargh argued that too many related words could make participants consciously aware of the theme and cancel an unconscious effect. Doyen et al. did find some evidence that participants could identify the elderly theme when directly probed — but nearly all denied noticing any connection between the sentence task and their walking.
Third, Bargh appealed to culture: priming works only if the association already exists in participants’ minds, so Belgian students might not link old age with slowness as American students do. Possible in principle — but Doyen et al. had adapted their materials to the Belgian sample by surveying 80 people about concepts associated with old age and selecting the frequent responses. Bargh et al.’s original article, by contrast, never measured whether their own participants associated the elderly stereotype with slower walking.
Finally, Bargh appealed to the accumulated literature: stereotype and behavioral priming had been replicated many times, so it was unreasonable to doubt the phenomenon over a single failure. This is persuasive only if the published record is an unbiased sample of all experiments run. That was precisely what the emerging replication crisis called into question. Journals favored novel, significant results and rarely published failed replications, so the sheer number of published successes could not reveal how often behavioral-priming experiments actually worked.
Doyen (2012) Was Not the First Failure
Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it, Commentators on the deleted blog post noted that Doyen et al. (2012) was not the first study to fail to reproduce elderly priming. As Nordbeck (2012) put it,
It is a bit of a shame that many of the arguments Bargh uses in his criticism of the Doyen study are arbitrary, unsupported and, on occasion, false in light of other research (even some from the area of priming) (Nordbeck, 2012).
The earlier failures had been published, but not as failures. They appeared as successful demonstrations of moderators: conditions under which the effect was supposed to grow, shrink, or reverse (Cesario et al., 2006; Dijksterhuis et al., 1998; Hull et al., 2002). Bargh later cited these very studies without noting that none of them had reproduced his basic result — a difference between an elderly prime and a control group.
That this could happen reflects a habit of the field. Social psychologists tracked patterns of statistical significance more than the magnitude of effects. Evidence for a moderator requires a significant interaction, and once an interaction turns up, attention shifts from the main effect to the conditional effects that explain it. The claim is no longer “priming works” but “priming works differently under condition X” — and a study can support that claim while quietly failing to reproduce the main effect it was built on.
Dijksterhuis et al. (1998) is the clearest case, appearing just two years after the original. Studies 2a and 2b each contained the two conditions needed to test elderly priming — an elderly prime and a neutral prime, followed by the same judgment task. In both, primed participants walked slightly more slowly, but the effects were small and nonsignificant, d = .25 and .29, against Bargh’s d = 1.04 and .79. Same direction, a quarter of the magnitude.
But testing Bargh was not the point of their paper; behavioral contrast was. Some primed participants also judged a specific elderly exemplar — the Dutch Queen Mother, Princess Juliana — and then walked faster than both the neutral and the ordinary elderly-prime groups. That contrast effect was significant in both studies and became the headline. The authors did notice that their ordinary elderly prime had failed to reproduce Bargh’s assimilation effect, called it “somewhat surprising,” and proposed that the intervening judgment task had wiped it out — reading the small same-direction means as “residual assimilation.” The failure was not overlooked; it was reinterpreted into a footnote. By this reading, the first evidence that elderly priming is not as large or robust as advertised appeared fourteen years before Doyen.
Their explanation is itself revealing. If inserting an innocuous judgment task can cut an effect by roughly 75%, then the effect is extraordinarily sensitive to procedural detail. Researchers wanted theoretically predicted moderators that would specify when and why priming occurs. What the evidence kept pointing to instead were unknown moderators: incidental features of a procedure that swing the effect from large to nothing. A phenomenon that surfaces only under a narrow and poorly understood set of conditions is not the robust effect Bargh’s original experiments described.
Dijksterhuis was only the first. Hull et al. (2002) ran the paradigm in two studies and found priming only among participants high in self-consciousness; the overall effects were moderate but nonsignificant, ds = .57 and .54, in small samples. Cesario et al. (2006) added a youth prime that sped participants up, but the elderly-versus-control difference was small and nonsignificant, d = .23. Jeffries and Fazio (2008) found an interaction with a stopping rule on an anagram task and no main effect at all, d = −.05. N. Wyer (2011) produced the first robust replication — d = .70 on walking speed and d = 1.03 on working memory. And in 2012, the same year as Doyen, one article reported strong effects in one study (ds = 1.26, 1.05) but weak ones in another (ds = .29, .31).
Seen in this company, Doyen’s failure is unremarkable. Some studies produce large estimates and others produce weak estimates and all estimates in small samples have large sampling error and make estimation of the true effect size impossible. The common solution to imprecise estimates in small studies is to combine the studies in a meta-analysis.
A Meta-Analysis of Elderly Priming
Existing meta-analyses of priming pool many kinds of primes and outcomes, and the resulting heterogeneity tells us little about any one paradigm. So I conducted a meta-analysis restricted to elderly priming: 15 tests from 12 studies in 8 articles. Because publication bias is evident in the broader literature, I used a selection model that estimates the bias and returns a corrected effect size (Vevea & Hedges, 1995), with clustered bootstrapping to handle multiple tests nested within articles and to build the confidence and prediction intervals.
The bias-corrected effect was d = .42, 95% CI [−.01, .71]. The point estimate is close to the much larger meta-analysis of Dai et al. (2023), but the interval is wide and includes zero. The studies are small and individually imprecise, and the heterogeneity is itself barely pinned down: tau = .26, 95% CI [.00, .37]. Nothing here compels the conclusion that Bargh’s and Doyen’s results reflect different true effects.
The estimated mean and variation of the population effect sizes produce a 95% prediction interval that ranges form d = -0.10 to d = 1.03.
Thus, a replication study so large that sampling error is negligible could land anywhere from essentially nothing to a very large effect, because the existing evidence simply does not fix the size of the effect.
The meta-analysis also cannot explain how Bargh got two significant results from samples of N = 30. If the true effect were d = .42, a study that small has only about a 20% chance of reaching significance, so two independent significant results would occur about 4% of the time. Either Bargh was lucky, or something in his particular setup inflated the effect — for instance, a confound in the word list, if “Florida” slows NYU students’ walking by evoking a laid-back-Southern stereotype rather than an elderly one.
In short, a meta-analysis of the studies available in 2012 would have shown Bargh that Doyen’s result was entirely consistent with the broader evidence, however much it clashed with his own. There was no need to reach for incompetence. His deeper error was assuming that his two original studies, together with conceptual replications using other primes, had already established a general effect of elderly priming on behavior.
The Fall-Out
Concerns about the credibility of social psychology were already building. Daniel Kahneman believed in priming and had featured it prominently in Thinking, Fast and Slow, yet he was growing worried about the field’s reputation and raised it with Bargh directly (Bartlett, 2012). In a widely circulated email to Bargh and other leading priming researchers, he proposed a simple remedy: if Doyen et al.’s failure was a fluke or a botch, the experts could settle the matter by running rigorous replications and showing the effects held. With an expected effect of d = .40, pooling resources for N = 400 would give roughly 98% power to detect it.
The collaboration never happened. Independent researchers ran preregistered replications instead, often with disappointing results — in Dai et al.’s (2023) meta-analysis, the preregistered studies, most of them replications, produced essentially no effect at all (d = .02).
In 2017, Kahneman returned to the problem, citing Overall (1969) on why underpowered research is not merely wasteful but pernicious: it inflates the share of false positives among published findings. Bargh’s experiments, with 15 participants per condition, had little power to detect any plausible effect and cannot account for the near-90% success rates in the published priming literature. Hundreds of studies manufactured the appearance of a robust phenomenon that low power, publication bias, heterogeneity, and near-zero preregistered effects do not support.
So the classic behavioral-priming paradigms simply faded from the research frontier. In that sense the field is practically dead — a stark case of strong belief outrunning weak evidence. Researchers can always find a reason to dismiss a failure and build ever more elaborate theory around inconsistent results, all without first establishing that the phenomenon the theory explains is real. As the meta-analysis above shows, even the 2012 evidence carried enough uncertainty to justify Kahneman’s call.
Which raises the obvious question: why wouldn’t a true believer, with the resources of a Yale professor, simply run the study and prove the point? Bartlett (2013) put it to Bargh directly:
So why not do an actual examination? Set up the same experiments again, with additional safeguards. It wouldn’t be terribly costly. No need for a grant to get undergraduates to unscramble sentences and stroll down a hallway.
Bargh was unenthusiastic. He wouldn’t ask graduate students worried about their job prospects to spend time on stigmatized research, and he was aware that some critics thought he had a “special touch” with priming — a word that sounds like praise but isn’t. “I don’t think anyone would believe me,” he said.
The same article also reports an interview with Axel Cleeremans, one of Doyen’s collaborators. Cleeremans reports that his collaborators and he found found the tone of Bargh’s letter insulting, but that they also took his criticism seriously enough to conduct another replication study that addressed his stated concerns. The study also failed to replicate the original results. Another unpublished replication failure was posted online by Pashler et al. (2011; Srivastava, 2012)
While the evidence that elderly priming works was mixed and inconclusive, Bargh expressed strong confidence in his findings and his theory of unconscious activation of behavior. In his blog post, he believed that the long-term future of priming research was secure because the phenomenon rested on a large and converging empirical foundation:
Like many scientists, I take the long view and continue to have faith in science as a cumulative process, with particular studies that vary in their methods and approaches converging on deeper underlying principles and mechanisms. No single experiment, standing by itself, should be the basis for concluding anything, maybe most especially in psychological research. And when a single study does not replicate another one whose findings are solidly embedded in theories of more than one scientific field and which is consistent with dozens if not hundreds of other conceptual replications, then responsible scientists—and responsible science journalists—do not rush to judgment and make claims that the entire phenomenon in question is illusory. Thomas Kuhn famously argued in The Structure of Scientific Revolutions that it is the accumulation of contrary evidence over time, not single studies, which is needed to overturn established concepts and principles in any branch of science.
A decade on, his faith in cumulative science has been vindicated — just not as he expected. Failure after failure, and no convincing new evidence from the original researchers, has left claims about unconscious behavioral priming no better established than Freud’s speculations about the unconscious. Priming researchers published and circulated their successes and buried their failures, manufacturing the look of robustness; and as shown here, even the successes never told a clean story or gave clear evidence that elderly priming had ever produced a reliable effect.
Bargh was right about one thing: no single failure should have settled it. What mattered was the accumulation of evidence. That accumulation just happened to vindicate Doyen’s skepticism rather than his own confidence. Maybe, somewhere in his unconscious, Bargh knows this too — which might be why he never tried to replicate his 1996 findings.
The Lesson
The Lesson
Bargh is not the first scientist to spend much of a career on an idea that never lived up to its promise. That is an occupational hazard of working at the edge of a field: the questions are hard, the methods imperfect, and a genuine breakthrough is difficult to tell from a false lead until many technical problems are solved. Strong belief supplies the motivation to push through those obstacles — but it also makes the warning signs of failed studies easier to miss. Economists call inflated asset prices “irrational exuberance”: a stock can stay overvalued for as long as enough investors share the illusion, then crash to its true worth. Psychological theories kept aloft by shared belief rather than evidence behave the same way, as the citation history of Bargh’s 1996 article shows.
That casual attitude toward evidence is on display in Matthew Lieberman’s 2012 blog post on Doyen’s failure — a post that mentions, in passing, an unpublished non-replication from Lieberman’s own lab:
There have been multiple unpublished non-replications of the elderly-walking (including one in my lab that I may discuss in a future blog), but it is hard to know what to make of them given that they haven’t been peer-reviewed.
The catch-22 is clean. We trust only peer-reviewed findings, but psychology journals publish overwhelmingly significant results (Sterling, 1959; Sterling et al., 1995). So the failures stay invisible — unless they are dressed up as significant interactions, or land in a venue like PLOS ONE that isn’t gatekept by the same community. More telling still is what Lieberman took the stakes to be:
At a certain level it does not matter whether the exact primes Bargh used produce a change in walking speed over the exact distance he measured. Some have said ‘We need to replicate this exactly. Conceptual replications aren’t good enough’. But I’m not sure why we care about this specific manipulation unless we are about to start using it as an intervention to treat patients. What we care about is whether priming-induced automatic behavior in general is a real phenomenon. Does priming a concept verbally cause us to act as if we embody the concept within ourselves? The answer to this question is a resounding yes.
The “resounding yes” rests entirely on a published record that reports successes and buries failures — and, worse, treats interaction effects that masked the very replication failures at issue as further confirmation. Shown only the hits, it is easy to conclude the phenomenon must be real. Kahneman (2017) later admitted he had trusted that record too much when he built on it in his book. The lasting contribution of the controversy is that younger psychologists now treat direct replication as something more than an optional courtesy.
Younger scientists can lower the risk by studying the previous generation’s mistakes. We are fond of saying that each generation stands on the shoulders of giants. We say less often that many equally gifted scientists climbed just as high and backed ideas that proved wrong. The ones whose theories survived were not necessarily wiser or more careful at the outset; to a real extent, nature happened to cooperate. Darwin’s mechanism of natural selection withstood ever more stringent tests; Lamarck’s mechanism of inheritance did not. Neither man could have known which way it would go when he set out.
Because you cannot know in advance whether nature will cooperate, the main lesson of what Kahneman (2012) called the “train wreck” of priming research is methodological before it is anything else: do not expect clean answers from small samples, and do not trust significant results from journals that never publish the null ones. And beyond method, a matter of temperament — stay skeptical, and resist the pull to become a believer. As Feynman put it:
The first principle is that you must not fool yourself, and you are the easiest person to fool.
Question: Is there credible evidence that automatic priming of behavior without people’s awareness works?
Current Scientific Answer:
There is no strong, broadly accepted evidence that subtle cues automatically and reliably change people’s everyday behavior without their awareness.
In the 1990s and early 2000s, psychologists published many studies suggesting that simply exposing people to certain words, images, or concepts could unconsciously influence behaviors such as eating, helping others, studying, or task performance. These findings attracted considerable attention because they implied that everyday behavior might be shaped by unnoticed environmental cues.
However, many of the most influential behavioral priming studies failed to replicate in later, larger, and more rigorous experiments. Others produced much smaller or less consistent effects than originally reported. As a result, confidence in the broad theory of automatic behavioral priming has declined substantially.
This does not mean that unconscious influences on behavior do not exist. People are clearly influenced by factors they are not fully aware of, and environmental features such as portion size, food availability, defaults, and convenience can reliably affect behavior. However, these effects are generally explained by mechanisms other than the classic idea that merely activating a concept (such as “achievement,” “elderly,” or “politeness”) automatically triggers corresponding behavior.
The current scientific consensus is therefore more cautious: there is no well-established, independently replicated demonstration showing that incidental, unnoticed conceptual primes reliably produce meaningful changes in everyday behavior. While some individual studies report positive effects, they have not accumulated into a robust, reproducible body of evidence comparable to well-established findings in other areas of psychology.
In short: If by automatic priming of behavior you mean that unnoticed words, images, or concepts reliably cause people to eat more, study harder, perform better, or otherwise change their everyday behavior, the best current answer is no. The stronger claims made two decades ago are not supported by the evidence available today.
The replication crisis has not been kind to priming research. Social priming has become the poster child for the kind of replication failures that are expected when researchers conduct underpowered studies, use flexible analyses, and selectively publish results that crossed, or nearly crossed, the conventional threshold for statistical significance, p < .05. The problem is not mysterious. If many small studies are run and only the successful ones enter the published literature, the published record will exaggerate effects, hide failures, and create the illusion of a robust phenomenon.
Yet defenders of social priming continue to insist that priming is real and robust. Their main defense is no longer a stream of new, well-powered, preregistered experiments showing that classic priming paradigms work. Instead, the defense has shifted to meta-analysis. The argument is that even if individual studies are weak, the combined literature still shows an effect. Dai et al. (2023) make this argument explicitly. They conclude that the field has moved beyond the question of whether priming exists and should now focus on mechanisms.
That conclusion is not supported by the evidence. A meta-analysis of p-hacked studies is not a cure for p-hacking. If the primary studies were produced by low power, selective reporting, flexible analysis, and publication bias, then the meta-analysis inherits those problems. A significant average effect does not show that priming is robust. It does not identify a reliable paradigm. It does not show that effects occur outside awareness. It does not explain why preregistered and direct replication attempts have often failed. And it does not justify moving from existence questions to mechanism questions.
An empirical science needs paradigms that can reliably produce the phenomenon under specified conditions. If behavioral priming is real and robust, proponents should be able to demonstrate it prospectively in well-powered studies. Cohen made this point long ago: science requires studies with a decent chance of detecting the effects they are designed to test. If the true effect is around d = .30 or d = .40, then the classic small priming studies were severely underpowered. Underpowered studies can generate significant results, but those results will be selected, inflated, and difficult to replicate.
This raises the obvious question: if priming works, why are its strongest defenders not still doing programmatic priming research? A robust phenomenon should generate active research programs: stable paradigms, preregistered replications, planned moderator tests, and cumulative theory development. Instead, much of the contemporary defense of priming relies on retrospective meta-analyses of old, heterogeneous, mostly nonpreregistered studies. Actions speak louder than words. If priming is robust, show that it works in the lab.
I have tried to publish a critique of Dai et al.’s meta-analysis in traditional psychology journals. Psychological Bulletin, which published the original meta-analysis, was not interested in publishing a correction of its conclusions. PSPR was also not interested in a paper that challenges the integrity of the evidentiary practices behind one of social psychology’s most famous literatures. Fortunately, journal gatekeeping does not prevent readers from evaluating the arguments for themselves.
The following list summarizes the main problems in Dai et al.’s meta-analysis and explains why their conclusion that behavioral priming is a robust phenomenon is not supported by their own evidence.
Scientific Problems in Dai et al. (2023) Meta-Analysis
A significant average effect is not evidence of a robust phenomenon.
Dai et al.’s main empirical result is a positive mean effect, d = 0.37, across 351 studies, 224 reports, and 862 effect sizes. They interpret this as evidence of a moderate and robust priming effect. But a positive grand mean only shows that the average coded effect is greater than zero. It does not show that any specific priming paradigm works reliably, that the effect is predictable, or that the literature has identified the conditions under which the effect occurs. This is the central inferential error. A heterogeneous collection of effects can have a positive mean even if many individual paradigms are unreliable, null, or false positive.
The literature is highly heterogeneous.
Dai et al. report substantial heterogeneity in the overall analysis and therefore use random-effects models. Their results show “substantial variability across studies and contexts,” but they still conclude that the effect is robust across procedures. That is too strong. High heterogeneity means the average effect is not a good description of individual studies or future studies. Robustness would require a narrow range of expected effects, or at least identifiable and replicated boundary conditions. Heterogeneity alone is not evidence for a robust phenomenon.
They do not report prediction intervals.
A random-effects meta-analysis with large heterogeneity needs prediction intervals. A confidence interval around the mean asks whether the average effect is different from zero. A prediction interval asks what effect size should be expected in a new study. That is the relevant quantity for “robustness.” If the prediction interval includes zero or negative effects, the literature does not support the claim that priming reliably produces positive behavioral effects. Your reanalysis indicates that Dai et al.’s estimates imply very wide prediction intervals, including negative effects, which directly contradicts the robustness claim.
Publication bias is present by their own analyses.
Dai et al. do not find a clean literature. They report evidence of small-study bias and funnel asymmetry. Their PET–PEESE analyses indicated that small studies reported larger effects, and the PEESE-adjusted estimate was reduced to d = 0.291, although still statistically significant. A reduced but still significant average does not eliminate the publication-bias problem. It merely says that one model still leaves a nonzero mean. It does not show that the published literature is robust, nor that individual paradigms are replicable.
Some bias-correction methods give much smaller estimates.
Dai et al. use multiple publication-bias adjustments, including PET–PEESE, Vevea and Woods selection models, and Mathur and VanderWeele sensitivity analyses. They acknowledge that selection models should be used to explore a range of estimates rather than to obtain one preferred estimate. The conclusion nevertheless emphasizes robustness. That is selective interpretation. A fair conclusion would stress the uncertainty and model dependence of the bias-corrected estimates, not reassure the field.
Their selection models do not handle dependence among effect sizes.
Dai et al. explicitly note that Vevea and Woods selection methods do not account for statistical dependence among effect sizes. This matters because their dataset contains multiple effect sizes from the same studies, samples, and reports. If selection-bias corrections are applied in a way that ignores dependence, the adjusted estimates and uncertainty intervals can be misleading. This weakens the evidentiary basis for strong claims about robustness.
The preregistered studies are essentially null.
This is one of the most damaging findings. Dai et al. report that only a small number of reports were preregistered, but those preregistered studies did not show the same effect pattern as the nonpreregistered literature. In your summary, the key contrast is preregistered studies around d = .02 versus nonpreregistered studies around d = .37. That contrast is fatal to a simple robustness claim. If priming were robust, preregistered studies should not collapse toward zero.
They dismiss preregistered failures asymmetrically.
Dai et al. explain the preregistered/null pattern by pointing to the small number of preregistered studies, replication difficulties, or possible differences between original studies and replications. That may be partly true, but it creates an asymmetric evidentiary standard. Positive nonpreregistered findings are treated as evidence for priming; preregistered failures are treated as inconclusive. This is precisely the kind of post hoc protection that prevents a theory from being falsified.
Most included studies were not preregistered.
Dai et al.’s descriptive table reports that among 224 included reports, only three reported preregistration. That means the positive average is overwhelmingly based on the older, flexible, nonpreregistered literature. This is exactly the kind of literature in which selective reporting, optional stopping, outcome switching, participant exclusions, and analytic flexibility can inflate effects.
Participant exclusions and covariate use are common.
Dai et al. report that many reports included covariates or excluded participants. Their study-quality section treats preregistration, covariates, and exclusions as indicators relevant to possible p-hacking. These features do not prove bias in any particular study, but they are red flags in a literature where the main evidentiary claim depends on many small, flexible, nonpreregistered experiments.
Their “quality assessment” is weak.
Dai et al. count random assignment, participant blindness, and signs of p-hacking. But random assignment was an inclusion criterion, so it cannot distinguish stronger from weaker studies. For blindness, only 37% of studies used funneled debriefing; for the rest, they state that it was hard to determine whether participants were blind, but that awareness was probably unlikely because priming effects are subtle. That is not a strong quality assessment. It assumes what needs to be shown.
Awareness is inadequately handled.
The controversial claim is not that cues can influence behavior when participants notice them. The controversial claim is influence without awareness of the prime’s relevance or influence. Yet Dai et al.’s own descriptive table shows that only 37.82% of effect sizes came from studies with funneled debriefing, and 84.22% used supraliminal primes. This means their broad meta-analysis does not primarily test the strongest implicit-priming claim.
Supraliminal cue effects are conflated with implicit priming.
Most studies in Dai et al. involve visible primes. Visible primes may influence behavior through ordinary cueing, demand, construal, motivation, or conscious goal activation. That is not the same as showing that behavior is guided outside awareness. Their conclusion about priming as a long-debated phenomenon slides between the broad, relatively uncontroversial claim that incidental cues can influence behavior and the stronger, controversial claim that behavior is automatically influenced outside awareness.
The dataset combines many theoretically distinct phenomena.
Dai et al. include primes involving achievement, common behaviors, money/marketing/finance, morality/God/prosociality, motivation, sex/gender/romantic behavior, and stereotypes. They also include verbal and visual primes, subliminal and supraliminal primes, performance outcomes, donations, choices, consumption, reaction times, and other behavioral measures. A positive average across this broad mixture does not validate a single phenomenon called behavioral priming. It may simply show that many different manipulations sometimes influence many different behaviors.
“Behavioral” versus “nonbehavioral” priming is too broad to identify mechanism.
Their core theoretical contrast is behavioral versus nonbehavioral primes. But they find no meaningful difference between the two: nonbehavioral primes d = 0.394 and behavioral primes d = 0.344, with B = 0.05, p = .14. They interpret this as suggesting overlapping processes. But a null difference between broad categories does not identify a mechanism. It may simply show that the categories are too crude, heterogeneous, or underpowered for mechanistic inference.
Moderator analyses are weak because the relevant moderator conditions are rare.
Dai et al. acknowledge that most included studies did not manipulate the key theoretical moderators: 89% did not manipulate goal value, 91% did not manipulate goal expectancy, and 82% did not include a filler task between priming and behavior. This makes strong moderator conclusions inappropriate. A literature cannot establish mechanisms if the relevant moderators are mostly absent.
Uneven moderator distributions reduce power.
Dai et al. explicitly state that the distribution of key moderators was highly uneven and that RVE has low power for moderators with uneven distributions. This is an important limitation. Moderator claims require interaction evidence, and interactions require much larger samples than main effects. A sparse and uneven moderator structure cannot support the claim that the field has moved from existence to mechanisms.
Simple effects are not evidence of moderation.
Even if some simple effects appear strong in certain moderator cells, the relevant evidence is the interaction. A significant effect in one condition and a nonsignificant effect in another is not evidence that the conditions differ. If the interaction evidence is weak, underpowered, or selected, moderator claims are unstable. This is especially important because the literature is already biased toward significant results.
Heterogeneity is treated as theoretical richness rather than an evidentiary problem.
Dai et al. repeatedly frame variability as consistent with different mechanisms or contextual dependence. But unexplained heterogeneity is not evidence for moderators. It is evidence that the effect is not yet predictable. The field cannot move to mechanism questions until it has reliable paradigms that produce the effect under specified conditions.
Their conclusion overstates what meta-analysis can show.
Dai et al.’s conclusion says their synthesis provides “solid evidence” that priming is real, not merely a bubble of publication bias, and that it should “reassure the field.” That is stronger than the data warrant. Meta-analysis of a selected, heterogeneous, mostly nonpreregistered literature can suggest that some effects may be real. It cannot establish that behavioral priming is robust as a general phenomenon.
The power implications are not drawn.
If the true effect is d = .29 to .37, typical small priming studies are underpowered. Many classic experiments with n = 20–40 per cell would have low power to detect effects of this size. Dai et al. do not make this implication central. If most original studies were underpowered, then significant findings are expected to be selected and inflated, and moderator analyses are even less credible.
The positive average is compatible with many false positives.
A significant average effect does not tell us how many individual significant findings are false positives. In a heterogeneous literature with low power and selection for significance, many significant findings can be false positives even if the mean effect is positive. This is why z-curve and related evidential-value analyses are needed. Dai et al.’s focus on mean effect size does not answer the false-discovery question.
Their methods estimate effects, not replicability.
Dai et al. assess mean effects and publication-bias-adjusted mean effects. But robustness is fundamentally about whether findings replicate. A field can have a nonzero average effect and still have poor replication prospects if studies are underpowered, selected, and heterogeneous. Their analysis does not directly estimate expected replication rates or discovery rates.
Failed replications are not integrated as severe tests.
Dai et al. cite failed replications as part of the motivation for the meta-analysis, but the conclusion effectively downgrades their importance once the average effect remains positive. That is not a balanced evidentiary treatment. Large, preregistered, or direct replication failures should receive greater weight as severe tests of the robustness claim, not be absorbed into a heterogeneous average and then explained away.
“Robust to publication bias” is too weak a criterion.
Their claim that the effect remains statistically significant after some bias adjustments does not establish robustness. A robust phenomenon should show reliable effects in well-powered, preregistered studies and should produce predictable effects in new studies. Surviving some bias corrections is a much lower standard.
The meta-analysis does not identify a robust paradigm.
The key scientific question is not whether the average across 862 effects is positive. The key question is: which priming paradigm reliably produces the effect? Dai et al. do not identify such a paradigm. Without a robust paradigm, the field cannot credibly study mechanisms, moderators, or applications.
One-off applied studies do not establish cumulative science.
The literature contains many isolated demonstrations involving different primes, outcomes, populations, and contexts. That pattern can produce a significant meta-analytic average while still lacking cumulative progress. A robust phenomenon should generate programmatic research around stable paradigms. Dai et al.’s synthesis does not show that this has happened.
The analysis treats conceptual replication as strength, but it also weakens interpretability.
A broad conceptual-replication database increases sample size, but it also combines many different true effects. That helps reject the null hypothesis that the average is zero, but it makes it harder to know what was actually established. If the procedures differ too much, the grand mean becomes theoretically thin.
The conclusion moves from “some effects may be real” to “the field should be reassured.”
The defensible conclusion is modest: some incidental priming effects may be real, and the average coded effect remains positive under some bias corrections. The stronger conclusion—that behavioral priming is robust and that the field can move beyond existence questions—is not supported.
The strongest claim about implicit influence is not established.
Priming with awareness is not controversial. The important claim is that incidental cues influence later behavior outside awareness of their influence. Dai et al.’s database is mostly supraliminal, often lacks funneled debriefing, and combines many visible cue manipulations. Therefore, even if their average effect were accepted, it would not establish the stronger claim that behavior is robustly influenced outside awareness.
Conclusion
Dai et al.’s conclusion that behavioral priming is robust is stronger than their evidence permits. Their own results show substantial heterogeneity, evidence of small-study and publication bias, sparse preregistration, null effects in preregistered studies, highly uneven moderator distributions, and limited assessment of awareness. A positive average effect across a heterogeneous and mostly nonpreregistered literature may show that some incidental cue effects are real, but it does not identify a reliable priming paradigm, establish robust boundary conditions, or justify moving from existence questions to mechanism questions. Robustness requires prospective replication under specified conditions, not merely a statistically significant grand mean.
Gary P. Latham is a professor at the business school of the University of Toronto, where he has spent his career studying goal pursuit — how people successfully accomplish goals. His first article appeared in 1973, “Effects of Goal Setting and Supervision on Worker Behavior in an Industrial Situation.” Over fifty years later he is still publishing (Budworth & Latham, 2025). He has authored around 200 articles and is among the most highly cited scholars in his field, with over 1,000 citations annually in recent years in Web of Science and more than 100,000 on Google Scholar. In short, Latham is a highly successful, highly cited scholar who remains deeply invested in his work.
While Latham has studied goal pursuit from many angles, one line of his research drew on Bargh’s automaticity model. Bargh proposed that behavior can be influenced by situational cues, or primes, without a person’s awareness. The classic study that seemed to demonstrate this showed undergraduates words related to the elderly and then found them walking more slowly afterward (Bargh et al., 1996). It took more than a decade for a replication failure to be published (Doyen et al., 2012) — which does not mean it was the first failed attempt. Colleagues had informally shared that they could not reproduce the finding, but such results were practically impossible to publish. So it was big news when a failure finally appeared in print.
Meanwhile, Daniel Kahneman — who had won the Nobel Memorial Prize in Economics — published a popular book that featured priming studies as strong evidence that behavior is shaped by stimuli outside awareness (Kahneman, 2011, Thinking, Fast and Slow). Kahneman later sent Bargh an open letter asking for new replication studies to establish that priming actually works. Bargh did not produce a convincing demonstration, and as other researchers ran their own replications, many failed.
During the replication crisis, while priming was faltering in basic social psychology, Latham was still running priming studies — and in his, it worked (Itzchakov & Latham, 2020). Social psychologists paid little attention to these successes in applied organizational research. I became aware of them only when I searched for strongly supported effects within a meta-analysis of over 800 priming results (Dai et al., 2023), which led me to examine organizational priming more closely.
Latham even wrote an article titled “The Effect of Priming Goals on Organizational-Related Behavior: My Transition from Skeptic to Believer.” He built priming into his theory of goal-directed behavior — even after serious doubts about the underlying phenomenon had taken hold in social psychology (Chen, Latham, & Itzchakov, 2021; Latham & Locke, 2018).
Bargh himself took notice, and welcomed the finding that priming appeared to produce strong effects in real-world settings:
“The research reviewed in Chen et al.’s (this issue) meta-analysis shows that a person’s goal pursuits and motivational states can be induced by external means, or ‘primes’, and then operate in much the same way as if the person made a conscious intention to pursue that goal. Previous demonstrations of this phenomena in psychology laboratories are now extended to real life organizations and settings and shown to produce even stronger effects than before” (Bargh, 2021).
These seemingly robust priming effects in organizational research piqued my curiosity, but I was skeptical — as Latham himself had been, before he turned from Saul to Paul. So I looked more closely at the evidence for priming effects on task performance. What could explain stronger, more robust results in organizational settings than in social psychologists’ own laboratories?
The Secret Sauce: Sample Size?
Priming is not a single, well-defined phenomenon. Studies vary along many dimensions — often called moderators — and even similar-looking studies can produce different results. Why did Bargh et al. (1996) find effects of elderly primes on walking speed across two studies, while Doyen et al. (2012) found nothing?
One obvious reason is chance. Every experiment is a gamble, because its outcome is shaped by sampling error — especially in small samples. A few disinterested participants can drag down the group primed to succeed; in the next study, that same kind of participant lands in the control condition and inflates the apparent effect. Small samples were the norm in priming research (Kahneman, 2017). Many studies were noisy gambles. What made this a crisis rather than mere noise is that the wins were published and the losses mostly were not (Shanks & Vadillo, 2021; Sterling et al., 1995).
To produce reliable evidence, either the effect has to be large or the sample has to be large enough to overcome chance. So the first plausible explanation for more robust organizational results is that researchers in this area — Latham among them — simply used larger samples. Call this the Sample Size as Moderator hypothesis.
We can test this against the most recent meta-analysis of organizational priming. The open dataset contains 69 effect-size estimates (Latham, Chen, Piccolo, & Itzchakov, 2023). Of these, 31 (45%) are statistically significant at the conventional α = .05 — barely above the 38% found across priming studies generally (Dai et al., 2023). A seven-point gap is no evidence that organizational priming is more robust; both literatures show priming failing about as often as it succeeds. Larger samples are not the secret sauce.
The Secret Sauce: Picture Priming?
Latham et al. (2023) aimed to include every study of priming in organizational settings. A large share came from Latham and his close collaborators — Shantz, Stajkovic, Itzchakov, Piccolo — as their co-authorships show. And when the dataset is broken down by lab, the pattern splits sharply. Studies from Latham and his collaborators succeed 78% of the time. Studies from every other lab succeeded only 24% of the time.
So it is not organizational settings that make priming robust. The effect is not a property of the field; it is a property of one group of researchers. What makes their studies different?
One possibility is the prime itself. Perhaps priming really is context-dependent, and Latham’s group happened onto a type of prime that works where others fail. There are a few candidates to consider. Some studies used subliminal words flashed briefly on screen — attractive because the stimulus is unambiguously outside awareness, but, at least in behavioral priming, apparently ineffective (Schimmack, 2026). Others embedded words in tasks such as scrambled sentences. Participants processed these consciously yet, in interviews, did not believe the words had influenced them — and in many cases they were right: the significant results were chance findings that did not replicate (Schimmack, 2026; Shanks et al., 2013; Shanks & Vadillo, 2019). Neither of these prime types survives scrutiny. But Latham’s signature manipulation was neither subliminal nor verbal. It was a picture.
In the seminal field study by Shantz and Latham (2009), 80 employees were randomly assigned to a control condition and a priming condition. The prime was a picture of a runner winning a race. The priming result was statistically significant, F(1, 77) = 4.95, but not remarkable. The effect-size estimate was moderate, d = .43, with a wide confidence interval ranging from .05 to .93. Notably, Shantz and Latham (2011) also obtained a significant result in a replication study with only a quarter of the sample size, total N = 20, t(18) = 2.23, p < .05. A second replication study with 44 participants was also significant, t(42) = 2.04, p < .05.
Studies with modest effect sizes and small samples are likely to produce non-significant results at some point (Schimmack, 2012). I asked professor Latham whether there were any unpublished studies with non-significant results. He replied that this was not the case.
“All my priming studies worked probably because I do pilot studies” —Latham, personal communication, July 30, 2026
Even without non-significant results, the published estimates probably benefit from inflation. With three independent studies the t-values should vary with a variance near 1. Here the variance of 2.23, 2.23, 2.04 is only 0.01. This is what the Test of Insufficient Variance (TIVA; Keiner & Renkewitz, 2019; Schimmack, 2015) is built to detect: a simulation with t-tests and matching degrees of freedom puts the probability of variance this small or smaller at about 1%. A representative set of studies would have produced some non-significant results — and, for these sample sizes, a lower average effect size.
The Secret Sauce: Feedback as Moderator
The studies with the strongest effects in Dai et al.’s meta-analysis came from Itzchakov and Latham (2020), which reported four studies. Importantly, the main hypothesis was not simply that primes influence behavior. It was that the effect of primes is moderated by performance feedback — primes should work more strongly when paired with feedback.
Supporting this requires a significant interaction, and all four studies delivered one. Once more, luck seems to have helped: the four interaction t-values cluster implausibly tightly, with a variance of 0.014. With samples near 200 the test statistics are effectively normal, so the analytic test suffices: a variance this small has probability p = .002.
More important are the results for the no-feedback condition that replicates the previous studies. Across the four no-feedback contrasts, three were null (d = 0.35, −0.14, −0.02) and one was significant (Study 4, d = 0.71). Pooled, the no-feedback effect is d = 0.13, with a 95% confidence interval spanning zero, [−0.10, 0.36]. The condition that most directly reproduces the original Shantz and Latham design — a picture prime and nothing else — shows essentially no effect once the four studies are combined.
In other words, Itzchakov and Latham had four opportunities to reproduce Latham’s foundational priming effect in the no-feedback conditions. Three produced nothing, one was significant, and pooled across all four the effect is indistinguishable from zero (d = 0.13, [−0.10, 0.36]). None of this was reported as a replication failure. The studies were presented as successes — because the interaction was significant — with the weak no-feedback effects framed merely as smaller than the feedback effects. But a picture prime that does essentially nothing on its own, and appears to work only alongside feedback, is not the phenomenon Latham built his theory on. It is a narrower claim: not that priming works and feedback amplifies it, but that priming may not work at all without feedback.
Whether priming genuinely works in combination with feedback is a question for future studies. There is a plausible mechanism — feedback could make a goal more concrete or more relevant to performance. But that is a far narrower claim than the original one: that a picture prime, by itself, improves work performance.
Conclusion
Latham described himself as a believer. Kahneman was a believer too. In Thinking, Fast and Slow, he asked readers to believe as well:
“Disbelief is not an option. The results are not made up, nor are they statistical flukes. You have no choice but to accept that the major conclusions of these studies are true.”
Years later, he recanted. In 2017 Kahneman wrote, “I placed too much faith in underpowered studies.” He had come to see that hundreds of statistically significant findings amount to strong evidence only if the studies are reported honestly. If there is a large file drawer of unpublished failures, the published literature can manufacture the illusion of a robust effect.
The evidence of publication bias in the priming literature is no longer in doubt (Dai et al., 2023). The harder question is whether any priming effect survives once that bias is taken into account. In organizational priming the answer appears to be: maybe, but only under narrow conditions. The evidence does not show that priming reliably changes behavior in general. It shows that some effects may appear in specific settings, with specific primes, specific outcomes, and sometimes only when paired with feedback.
Priming is not the first case of scientists finding compelling evidence for something that was not there. Langmuir called it pathological science: honest researchers, following the ordinary rules of their field, converging on an effect that does not exist (Langmuir, 1953). What makes it pathological is not fraud but conviction. The scientists most likely to fool themselves are the ones most committed to the phenomenon — the believers, driven to demonstrate what they already know to be true. The skeptic who doubts everything rarely produces a striking result; the believer who doubts nothing produces a career’s worth. Feynman put the hazard plainly: “The first principle is that you must not fool yourself — and you are the easiest person to fool” (Feynman, 1974).
Latham called himself a believer. That was the problem.
References
Bargh, J. A. (2021). Unconscious goal pursuit in real-life organizations: Commentary on Chen, Latham, Piccolo, and Itzchakov (2020). Applied Psychology: An International Review, 70, 254–261. https://doi.org/10.1111/apps.12259
Bargh, J. A., Chen, M., & Burrows, L. (1996). Automaticity of social behavior: Direct effects of trait construct and stereotype activation on action. Journal of Personality and Social Psychology, 71(2), 230–244. https://doi.org/10.1037/0022-3514.71.2.230
Chen, X., Latham, G. P., Piccolo, R. F., & Itzchakov, G. (2021). An enumerative review and a meta-analysis of primed goal effects on organizational behavior. Applied Psychology: An International Review, 70, 216–253. https://doi.org/10.1111/apps.12239
Dai, W., Yang, T., White, B. X., Palmer, R., Sanders, E. K., McDonald, J. A., Leung, M., & Albarracín, D. (2023). Priming behavior: A meta-analysis of the effects of behavioral and nonbehavioral primes on overt behavioral outcomes. Psychological Bulletin, 149(1–2), 67–98. https://doi.org/10.1037/bul0000374
Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral priming: It’s all in the mind, but whose mind? PLoS ONE, 7(1), e29081. https://doi.org/10.1371/journal.pone.0029081
Feynman, R. P. (1974). Cargo cult science. Engineering and Science, 37(7), 10–13.
Itzchakov, G., & Latham, G. P. (2020). The moderating effect of performance feedback and the mediating effect of self-set goals on the primed goal–performance relationship. Applied Psychology: An International Review, 69(2), 379–414. https://doi.org/10.1111/apps.12176
Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.
Langmuir, I. (1989). Pathological science (R. N. Hall, Ed.). Physics Today, 42(10), 36–48. (Original work presented 1953). https://doi.org/10.1063/1.881205
Latham, G. P. (2018). The effect of priming goals on organizational-related behavior: My transition from skeptic to believer. In G. Oettingen, A. T. Sevincer, & P. M. Gollwitzer (Eds.), The psychology of thinking about the future (pp. 392–404). Guilford Press.
Latham, G. P., Chen, X., Piccolo, R. F., & Itzchakov, G. (2023). An updated meta-analysis of the primed goal–organizational behaviour relationship. Royal Society Open Science, 10(4), 221494. https://doi.org/10.1098/rsos.221494
Latham, G. P., & Locke, E. A. (2018). Goal setting theory: Controversies and resolutions. In D. S. Ones, N. Anderson, C. Viswesvaran, & H. K. Sinangil (Eds.), The SAGE handbook of industrial, work & organizational psychology (2nd ed., Vol. 1, pp. 103–124). Sage.
Renkewitz, F., & Keiner, M. (2019). How to detect publication bias in psychological research: A comparative evaluation of six statistical methods. Zeitschrift für Psychologie, 227(4), 261–279. https://doi.org/10.1027/2151-2604/a000386
Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17(4), 551–566. https://doi.org/10.1037/a0029487
Schimmack, U. (2026). A new look at implicit priming: Making sense of heterogeneity in conceptual replication studies [Manuscript submitted for publication]. Department of Psychology, University of Toronto.
Shanks, D. R., Newell, B. R., Lee, E. H., Balakrishnan, D., Ekelund, L., Cenac, Z., Kavvadia, F., & Moore, C. (2013). Priming intelligent behavior: An elusive phenomenon. PLoS ONE, 8(4), e56515. https://doi.org/10.1371/journal.pone.0056515
Shanks, D. R., & Vadillo, M. A. (2021). Publication bias and low power in field studies on goal priming. Royal Society Open Science, 8(10), 210544. https://doi.org/10.1098/rsos.210544
Shantz, A., & Latham, G. P. (2009). An exploratory field experiment of the effect of subconscious and conscious goals on employee performance. Organizational Behavior and Human Decision Processes, 109(1), 9–17. https://doi.org/10.1016/j.obhdp.2009.01.001
Shantz, A., & Latham, G. P. (2011). The effect of primed goals on employee performance: Implications for human resource management. Human Resource Management, 50(2), 289–299. https://doi.org/10.1002/hrm.20418
Sterling, T. D., Rosenbaum, W. L., & Weinkam, J. J. (1995). Publication decisions revisited: The effect of the outcome of statistical tests on the decision to publish and vice versa. The American Statistician, 49(1), 108–112. https://doi.org/10.1080/00031305.1995.10476125
Vadillo, M. A., Hardwicke, T. E., & Shanks, D. R. (2016). Selection bias, vote counting, and money-priming effects: A comment on Rohrer, Pashler, and Harris (2015) and Vohs (2015). Journal of Experimental Psychology: General, 145(5), 655–663. https://doi.org/10.1037/xge0000157
We can distinguish two types of meta-analysis. Meta-analyses of direct replication studies combine studies with the same or very similar population effect sizes. So-called fixed-effect meta-analyses produce more precise estimates of the shared population effect size, because the sampling error of the combined data is smaller than the sampling error of the individual studies.
However, psychologists often conduct conceptual replication studies that vary conditions, stimuli, and dependent variables. Meta-analyses of these studies are more challenging because the studies have varying population effect sizes (van Erp et al., 2017). As a result, the average effect size is relatively uninformative and provides no information about the effect sizes of specific studies. The standard solution has been to model this heterogeneity by assuming a normal distribution for the population effect sizes, in addition to the normal distribution of sampling error. The previous post showed that (a) this assumption can lead to biased estimates when the true distribution of population effect sizes is not normal, and (b) developments in statistics make the assumption unnecessary. Meta-analysts can therefore drop the problematic normality assumption by adopting more flexible mixture models that make no assumption about the shape of the population effect-size distribution.
What is a Positive Effect?
The main contribution of this post is to reveal a theoretical contradiction between the assumption of normally distributed population effect sizes and the way studies are coded in meta-analyses. To see the problem, it helps to ask a simple question: what is a positive effect?
The difficulty of defining a positive outcome is well recognized in medicine. A positive diagnosis usually means bad news — except when people hoping for a child learn that they are pregnant. Cochrane meta-analyses address this by coding results not only by the direction of change in an outcome (increase or decrease) but also by whether that change is beneficial or harmful: for mortality, an increase is harmful; for longevity, an increase is beneficial. For the meta-analyst, what matters is keeping the sign of an effect constant across studies, while the interpretation depends on the substantive meaning of the average. “Women live longer than men” is a positive difference if women are coded 1 and men 0, and a negative difference if the coding is reversed.
The default coding in a meta-analysis that tests a prediction is to assign signs according to the predicted direction: results supporting the hypothesis receive a positive value, and results contradicting it receive a negative sign.
This coding, combined with the assumption of a normal distribution of population effect sizes, leads to a paradoxical implication. The paradox is clearest in the extreme case where the average effect size is zero. Under the normality assumption, a mean of zero implies that half of the true effects are consistent with the prediction and half are opposite to it — that the theory predicts the direction of the effect no better than a coin flip. This implication is usually implausible, which is better read as evidence against the distributional assumption than as a verdict on the theory. Even when the average effect is small but non-zero, normally distributed heterogeneity implies that a substantial proportion of true effects run opposite to the prediction.
This state of affairs is often overlooked when meta-analytic results are interpreted. For example, Chen et al. (2025) reported large heterogeneity, with a 95% prediction interval from −1.05 to 1.77, implying that many studies have large true effects opposite to the predicted direction. Yet none of the published studies reported such results. The prediction interval implies a population of wrong-direction true effects that does not appear in the published record, and this implication went unexamined: are these effects real but suppressed by publication bias, or are they phantoms produced by an incorrect distributional assumption?
In short, coding a meta-analysis by agreement with a theoretical prediction, combined with the assumption of a normal distribution of effect sizes, predicts that many studies have true effects in the wrong direction. Because the actual data often show no such results, the prediction raises a question: do these wrong-direction effects exist and were suppressed by publication bias, or are they phantom effects manufactured by a false distributional assumption?
Modeling Only Positive Effect Sizes
When the data contain few negative effect-size estimates, the normality assumption can do more than inflate estimates of heterogeneity; it can also bias the effect-size estimate itself. The reason is that sampling error produces many negative estimates for small effects, even when the population effect is positive. If the data do not show these expected negative results, a model that assumes complete data infers that the positive results must have been produced by a larger population effect. This inflates the effect-size estimate, especially when the average effect is small.
To avoid this inflation, the truncation of the data at zero must be modeled. For example, if the true population effects follow a normal distribution centered at zero, a model that knows the data are truncated at zero expects a half-normal distribution and correctly infers that the mean of the full distribution is zero. A model that does not know the data are truncated fits a full normal distribution to the observed half-normal data and estimates a mean greater than zero.
To illustrate this, I ran a simulation with 64 conditions, varying the mean of the population effect sizes (4 levels), their standard deviation (4 levels), and the proportion of true null hypotheses (4 levels) — the same design used previously to show the problem of assuming normality when the true distribution is non-normal. The previous simulation showed that models with flexible distribution assumptions perform better than models that assume a normal distribution when the actual distribution is not normal.
This time, the simulation selected only positive effect-size estimates. The prediction was that ashr and RMA would be biased, because they assume no selection on the sign of an effect, whereas z-curve and weightr are selection models that can be specified to model selection on sign. Weightr, however, may still be biased when the distributional assumption is violated, because its selection-weight parameters can absorb the non-normality, fitting a non-normal distribution by distorting the estimated selection process. Z-curve was expected to perform well because it makes no assumption about the shape of the effect-size distribution and can model truncated data.
The results confirmed these predictions, which follow directly from the models’ assumptions. Z-curve had the smallest error (RMSE = .046), followed by RMA and ashr (RMSE = .094), with weightr the worst (RMSE = .168).
The poor performance of the weight-function selection model is noteworthy, because it is commonly assumed to produce more credible results when data are selected. This confidence is justified only when the distributional assumption holds. When it does not, the selection model can produce worse estimates than a model that does not correct for selection at all, such as RMA. A false distributional assumption can even yield negative estimates of the average population effect when all studies have positive population effects.
Finally, it is worth noting that ashr was not designed for meta-analysis, but to analyze data from studies with many predictor variables, where the sign of an effect is often arbitrary and roughly equal numbers of positive and negative results are expected. When all of these results are available, as in genomics, the model works well — indeed, without selection it is equivalent to z-curve. The bias identified here arises only when a model built for complete data is applied to selected meta-analytic results (e.g., van Zwet et al., 2024).
In conclusion, for meta-analysis it is problematic to assume that negative effect sizes estimates are common and produced by studies with negative population effect sizes. It is therefore necessary to allow for selection based on the sign of an effect size independent of selection for statistical significance. The most suitable models for meta-analyses therefore need to assume selection for sign without assuming a particular distribution of population effect sizes. The only model that checks both boxes is z-curve.
Illustration with Actual Data
Chen et al.’s (2025) meta-analysis of terror management studies serves as a useful example, because it analyzed the data with both a random-effects model and a weight-function selection model, and because the set of studies is large (k = [FILL IN]). As noted earlier, the models that assumed normality estimated very high heterogeneity, with a prediction interval extending well into negative values — implying that a substantial proportion of studies have true effects opposite to the theoretical prediction.
The histogram shows the distribution of the observed effect-size estimates. There are many small positive estimates, but negative estimates are rare. This is not what a normal distribution would produce: a normal distribution centered on the positive mean, with the estimated heterogeneity, would place substantial mass below zero and thus predict many negative estimates. The near-absence of negative estimates therefore points to one of two conclusions — either negative results were suppressed by publication bias, or the true distribution of effects is not normal.
Chen et al. (2025) did not model selection for sign, despite the visible asymmetry around zero. I refitted the model with a selection step at p = .5 (one-tailed), distinguishing positive (p < .5) from negative (p > .5) results. This model estimated a negative average effect, d = −.336, with large heterogeneity (tau = .956). A negative mean for a literature whose observed effects are overwhelmingly positive is implausible on its face. It arises because the model can only reconcile the near-absence of negative results with a normal distribution by assuming that a large number of negative and null results were suppressed — so many that the true distribution is centered below zero, and the observed positive results are merely its selected upper tail. This is possible only if one accepts both that negative results were heavily suppressed and that the true effect distribution is normal.
The random-effects model faces a different problem. It assumes no selection and fits a normal distribution to a set of effect sizes that are visibly skewed. Fitting a symmetric normal to this right-skewed distribution inflates the mean, which RMA estimates at .81.
The ash model makes no assumption about the shape of the effect-size distribution and produces a similar mean, .80, but a smaller median, .62, reflecting the right skew visible in the observed data.
Both RMA and ash, however, take the observed data at face value, even though the data contain few negative results. Z-curve can model truncated data without assuming a normal distribution. The next figure shows the z-curve plot with local power and local effect-size estimates. This model does not assume publication bias.
Z-curve models truncation of the data at zero and makes no distributional assumption about the population effect sizes. The z-curve plot shows the histogram with the fitted model, along with local average effect sizes below the x-axis and local power estimates beneath them. Even the non-significant results are estimated to have moderate effect sizes (~.5), and the overall mean is .77.
As with ash, the median is smaller than the mean — .63 versus .77 — reflecting the same right skew. That two flexible methods independently recover both a mean near .8 and a median near .62 indicates the skew is a real feature of the effect distribution, not an artifact of either method. A random-effects model cannot represent this asymmetry, because assuming normality forces its mean and median to coincide.
Across models, then, the mean estimates agree reasonably well — z-curve (.77), ash (.80), and the random-effects model (.81) all indicate a moderate-to-large average effect — with the weight-function selection model the lone exception, producing an implausible negative estimate. For this literature, truncation and the normality assumption have a relatively small effect on the point estimate, consistent with the simulations, which showed close agreement among models when the average effect is large. This is the favorable case, however: when the average effect is small, the normality assumption’s implied negative effects and the truncation at zero have a much larger impact.
So far, all models assumed no selection for statistical significance. This is unlikely, because published articles predominantly report significant results that support theoretical predictions (Sterling et al., 1995), and Chen et al. (2025) found evidence of publication bias in this literature. Z-curve can model publication bias by fitting only the significant results and estimating the distribution of the non-significant results that were not reported.
Taking publication bias into account has two effects. First, the effect-size estimates decrease. Second, the average is now computed over a larger reference set that includes the estimated non-significant studies, which have smaller effects. The mean effect size is reduced to .52 and the median to .34. Even after this correction the distribution remains right-skewed — the mean exceeds the median — and the median of .34 is the more representative value for a typical study.
In conclusion, meta-analytic models that make different assumptions produce different results. Making sense of these differences requires testing the assumptions and preferring models that avoid assumptions that cannot be tested. In this example, the data show clear evidence of selection on the sign of the effect, further evidence of selection against non-significant positive results, and evidence that the population distribution is not normal (median < mean). A model that accommodates these features of the data — z-curve — produced lower estimates than the random-effects model, while the weight-function selection model produced an implausible negative estimate.
It is not possible to generalize from this single example to all meta-analyses, but the results raise the possibility that many published effect-size estimates are biased by false distribution assumptions, especially when publication bias is also present. To assess the extent of these biases, such literatures should be reexamined with methods that can model selection and avoid the untestable assumption of normally distributed population effects — as z-curve does.
Correction (2026-07-14) The original post claimed that violations of the normality assumption biases standard random effects estimates. That finding was based on a coding error in the simulation program. The revised post shows the results for simulations with normal distributed effect sizes for H1 and mixtures of H0 and H1. In these simulations only estimates of the weight-function selection model (weightr package) are biased when the unimodal normal distribution assumption is violated.
Introduction
The main purpose of meta-analysis is to combine the results of quantitative studies. The simplest form averages the effect-size estimates of studies. The average is a better estimate of the population effect size because it draws on the combined sample size of the individual studies, and larger samples have less sampling error. A more sophisticated version takes each study’s sampling error into account, weighting studies with larger samples more heavily than those with smaller samples. This weighted average approximates the estimate one would obtain by pooling the raw data, under the assumption that all studies estimate the same effect.
These so-called fixed-effect meta-analyses assume that the samples are all drawn from the same population and that studies used interchangeable procedures, so that they all estimate a single population effect size. This assumption is defensible for a few phenomena grounded in basic sensory or cognitive architecture. Weber’s law, for instance — that the just-noticeable difference between two stimuli is a roughly constant proportion of their magnitude — reflects a property of sensory transduction and varies little across procedures and individuals. But such near-constants are the exception. Most effects in psychology vary with situational and personality factors, within and across cultures. This variation adds real differences in the true effect sizes across studies, over and above sampling error.
In recognition of this true variation, statisticians developed random-effects models (DerSimonian & Laird, 1986). To model the variation in population effect sizes, it is necessary to select a model that describes the unobserved variation in population effect sizes. From a statistical perspective, the easiest model assumes that population effect sizes have a normal distribution. The symmetry of this assumption implies that the estimate obtained from a set of studies is an unbiased estimate of the true average because positive and negative deviations from the average population effect size cancel each other out, just like random sampling errors cancel each other out.
Distribution Assumptions
From a theoretical perspective, it is unlikely that population effect sizes are normally distributed, particularly when the mean effect is small relative to the between-study heterogeneity. The reason is that effects are typically coded so that positive values are consistent with a directional theoretical prediction (e.g., conflicting colors produce slower responses in the Stroop task). While the size of the effect may vary across studies, it is often implausible that the true effect in some studies would be reversed (that conflicting colors would genuinely speed responses). A normal distribution, however, has support over the entire range of effect sizes and therefore always assigns some probability to negative true effects — an implication that is difficult to justify for a directionally predicted effect. When the mean is small relative to the heterogeneity, a normal distribution implies that a non-trivial proportion of true effects are negative, which is often theoretically implausible.
In practice, however, when heterogeneity is small, violations of the normality assumption have small and often negligible effects on meta-analytic estimates of the average effect (Hedges & Vevea, 1996). This helps explain why widely used random-effects models share the assumption that population effect sizes are normally distributed, including standard random-effects meta-analysis (DerSimonian & Laird, 1986) as implemented in commonly used software (Viechtbauer, 2010), selection-model approaches (Vevea & Hedges, 1995; Vevea & Woods, 2005), and the random-effects components of Bayesian model-averaging methods (Bartoš et al., 2023; Maier et al., 2023).
Necessary and Unnecessary Assumptions
It is useful to distinguish between necessary and unnecessary assumptions. All statistical methods require some assumptions to reduce the complex information in data to an interpretable statistic. However, not all assumptions of a model are necessary. In general, models with fewer unnecessary assumptions are preferable because they are robust to violations of these unnecessary assumptions by design. For example, the Pearson correlation coefficient assumes normal distribution of the two variables, whereas rank-order correlations do not. This makes rank-order correlations robust to violations of the normal distribution assumption and it is common practice to prefer rank-order correlations when variables’ distributions are not normal.
In meta-analyses, the assumption that population effect sizes are normally distributed is an unnecessary assumption. The idea of estimating the distribution of effects without assuming a parametric form is not new. Laird (1978) and Lindsay (1983) showed that the mixing distribution in a mixture model can be estimated nonparametrically by maximum likelihood, and that the resulting estimate takes the form of a discrete distribution on a finite set of support points. This provided, in principle, a fully flexible way to model heterogeneity without assuming normality. In practice, however, fitting these models was computationally demanding at the time.
Advances in optimization and computing have since removed these hurdles. Efficient algorithms for estimating mixture models — including modern convex-optimization methods (Koenker & Mizera, 2014; Kim et al., 2020) — now make it possible to fit flexible mixtures quickly and reliably. These developments led to the widespread adoption of nonparametric and empirical-Bayes mixture models in fields that analyze large numbers of effects, most notably genomics (Efron, 2010; Stephens, 2017), where the goal is to characterize the distribution of effects across thousands of tests and to identify which are likely to be real. Similar approaches have been applied in astronomy, education, and other areas that deal with many parallel estimates.
Heterogeneity of Population Effect Sizes
Meta-analysis, however, did not adopt these models, and continued to rely on the normal random-effects model. One reason is that the goal of most meta-analyses has been to estimate a single average effect, for which the shape of the effect-size distribution is largely irrelevant. The additional information provided by a flexible mixture — the structure of the heterogeneity itself — was not part of the question being asked.
This has changed with growing recognition of the distinction between direct and conceptual replication (Zwaan et al., 2018). Experimental psychologists rarely repeat a procedure exactly. Instead, they typically vary paradigms across studies in the belief that this is good scientific practice — probing moderators, boundary conditions, and the generalizability of an effect. As a result, the studies combined in a meta-analysis often estimate genuinely different population effects, and meta-analyses of such conceptual replications frequently show high estimates of heterogeneity (van Erp et al., 2017). In these cases the average effect size is of relatively minor interest, because it merely centers a wide distribution of different population effect sizes — an average across paradigm variants that no single study actually instantiates. The scientifically interesting question is not the average, but how and why the effects differ.
Mixture Models
Mixture models have another advantage over random effects models for meta-analysis. In standard meta-analytic models the population distribution of effect sizes is characterized by the mean and standard deviation of a normal distribution. Importantly, this information is removed from the observed data. It is therefore not possible to use this information to make sense of the heterogeneity in the observed effect size estimates in the data. A large effect size in the data may correspond to a small population effect size because it was inflated by sampling error and vice versa. This explains why heterogeneity often plays a minor role in the interpretation of meta-analyses. Heterogeneity exists, but it cannot be explained.
In contrast, mixture models combined with statistical tools that correct for regression to the mean make it possible to obtain estimates of the population effect size of individual studies. Often these effect size estimates for single studies have large uncertainty (wide confidence intervals), but they can be averaged to identify subgroups of studies with large effect sizes. This makes it possible to explore the heterogeneity in observed studies.
Simulation Study
To demonstrate the capabilities and advantages of mixture models over traditional random-effects models, I conducted a simulation study. To isolate the effect of the distributional assumption, the simulation used unbiased data (no publication selection). This removes selection as a confound and allows inclusion of the adaptive-shrinkage mixture model ash (Stephens, 2017), which was developed for genomics and assumes complete, unselected data. The other two models — the weight-function selection model (Vevea & Woods, 2005) and z-curve — model publication bias but reduce to unbiased estimation when no selection is present.
Z-curve is a meta-analytic mixture model (Brunner & Schimmack, 2020; Bartoš & Schimmack, 2022). Version 1 used the noncentrality parameters (ncp) of significant results to estimate the expected replication rate for direct replication studies. Version 2 used the ncp distribution across all studies to estimate the expected discovery rate. Version 3 adds an empirical-Bayes function that estimates population effect sizes by multiplying ncp estimates by their corresponding sampling errors, recovering effects from the latent distribution of noncentrality parameters.
The simulation used fixed sample sizes across studies, which ensures that mixture models fit in the effect-size metric and in the ncp metric produce identical results. Each condition contained k = 1,000 studies; the large k yields small sampling error in the estimates, making systematic biases more visible. The design crossed 4 means and 4 standard deviations of the population effect-size distribution with 4 proportions of true null hypotheses, producing 64 conditions. The population effect sizes when H1 is true were drawn from a normal distribution. However, mixing true H0 and H1 produces non-normal, bimodal distributions. This violates the distribution assumption of standard meta-analytic methods and could produce biased estimates.
These results for the mixture models (z-curve, ash) are not surprising. The models make no distribution assumptions and estimates of the mean closely match the simulated true values (z-curve RMSE = .009; ash RMSE = .010). The weight-function selection model had worse fit (RMSE = .138), but the random effects model fit as well as the mixture models (RMA RMSE = .003).
The results for the estimate of heterogeneity (tau) showed that all models estimated heterogeneity well, (z-curve RMSE = .018, ash RMSE = .027, weightr RMSE = .030, RMA RMSE = .013).
The present results show that the random effects model is robust to violations of the normality assumption even when heterogeneity is large and distributions are bimodal. The same cannot be said for the selection model with a normal assumption (weightr). Here the weight parameters to model selection bias can be biased to fit a normal distribution to non-normal observed data. The main implication is that it is possible to improve the performance of selection models by avoiding the normality assumption and modeling the distribution of population effect sizes with a flexible mixture model.
Implications
The present simulation showed that the random-effects model estimates the mean effect well when data are unbiased, even when its distributional assumptions are violated. This is consistent with a substantial prior literature. Simulation studies have repeatedly found that the pooled mean from the normal random-effects model is robust to non-normal random-effects distributions, including skewed and mixture distributions (Yamaguchi et al., 2017; Rubio-Aparicio et al., 2023; Baur et al., 2025). The robustness follows from the estimator itself: the pooled mean is an inverse-variance weighted average, which is consistent for the population mean regardless of the shape of the effect distribution. Large numbers of studies further protect against non-normality (Rubio-Aparicio et al., 2023), and with k = 1,000 the present simulation is in this favorable regime.
Where the normality assumption does matter is not the mean but the characterization of the distribution around it. The same literature finds that non-normality degrades the coverage of confidence and prediction intervals far more than it biases the mean (Baur et al., 2025), and that a symmetric normal model produces a wrongly symmetric prediction interval and a misleading heterogeneity summary when the true distribution is skewed (Yamaguchi et al., 2017). The recognized advantage of flexible mixture and semi-parametric models is therefore not lower bias on the mean but their ability to reveal latent structure — clusters and subgroups of studies — that the mean and a single heterogeneity parameter discard (Baur et al., 2025).
These are real but secondary limitations. The more fundamental problem is one that no distributional flexibility can address: the random-effects model, and the adaptive-shrinkage mixture model (ash) alike, assume that the observed studies are an unbiased sample of the population. In many literatures this assumption is untenable. In psychology, studies with significant results are far more likely to be published; the proportion of significant results often exceeds 90% (Sterling et al., 1995). Such publication bias inflates effect-size estimates, and correcting for it requires modeling the selection process — something neither the random-effects model nor ash attempts. A flexible distribution alone is not enough; what is needed is a model that combines distributional flexibility with a model of selection.
Z-curve occupies this niche: it combines a selection model, which corrects for publication bias, with a flexible mixture model, which captures heterogeneity without a distributional assumption. This is why z-curve outperforms weight-function models when studies are both heterogeneous and affected by publication bias (Schimmack, 2026).
References:
Bartoš, F., Maier, M., Wagenmakers, E.-J., Doucouliagos, H., & Stanley, T. D. (2023). Robust Bayesian meta-analysis: Model-averaging across complementary publication bias adjustment methods. Research Synthesis Methods, 14(1), 99–116.
Böhning, D. (2000). Computer-assisted analysis of mixtures and applications: Meta-analysis, disease mapping and others. Chapman & Hall/CRC.
DerSimonian, R., & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials, 7(3), 177–188.
Hedges, L. V., & Vevea, J. L. (1996). Estimating effect size under publication bias: Small sample properties and robustness of a random effects selection model. Journal of Educational and Behavioral Statistics, 21(4), 299–332.
Laird, N. (1978). Nonparametric maximum likelihood estimation of a mixing distribution. Journal of the American Statistical Association, 73(364), 805–811.
Lindsay, B. G. (1983). The geometry of mixture likelihoods: A general theory. The Annals of Statistics, 11(1), 86–94.
Maier, M., Bartoš, F., & Wagenmakers, E.-J. (2023). Robust Bayesian meta-analysis: Addressing publication bias with model-averaging. Psychological Methods, 28(1), 107–122.
Stephens, M. (2017). False discovery rates: A new deal. Biostatistics, 18(2), 275–294.
Vevea, J. L., & Hedges, L. V. (1995). A general linear model for estimating effect size in the presence of publication bias. Psychometrika, 60(3), 419–435.
Vevea, J. L., & Woods, C. M. (2005). Publication bias in research synthesis: Sensitivity analysis using a priori weight functions. Psychological Methods, 10(4), 428–443.
Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48.
Efron, B. (2010). Large-scale inference: Empirical Bayes methods for estimation, testing, and prediction. Cambridge University Press.
Kim, Y., Carbonetto, P., Stephens, M., & Koenker, R. (2020). A fast algorithm for maximum likelihood estimation of mixture proportions using sequential quadratic programming. Journal of Computational and Graphical Statistics, 29(2), 261–273.
Koenker, R., & Mizera, I. (2014). Convex optimization, shape constraints, compound decisions, and empirical Bayes rules. Journal of the American Statistical Association, 109(506), 674–685.
Laird, N. (1978). Nonparametric maximum likelihood estimation of a mixing distribution. Journal of the American Statistical Association, 73(364), 805–811.
Lindsay, B. G. (1983). The geometry of mixture likelihoods: A general theory. The Annals of Statistics, 11(1), 86–94.
For method folks, the picture tells the full story: z-curve can estimate the true mean of a set of heterogeneous studies better than the weight-function model because the weight function model makes unrealistic assumptions about the distribution of population effect sizes. Added bonus: z-curves estimates are related to actual studies, whereas the estimates of weightr are population estimates that are not connected to the actual studies.
The root cause of the crises in psychology is poor training in scientific thinking and scientific methods. Period! I know because I have been teaching at a top-ranked university in North America for over 25 years now. The most common criticism in student evaluations is that my courses are not psychology courses, but statistics course. The reason: I use numbers when I present research findings. But most students can get a degree in psychology without using numbers. Graduate education does not help because students learn from a mentor, who also never learned to think quantitatively. So, psychology is the worst of both worlds. It is neither qualitative research that pays attention to people’s thoughts, feelings, or actual behaviors, nor is it a quantitative science that use valid quantitative information for the same purpose. It is a pseudo-science that produces meaningless numbers that mainly serve the purpose of claiming scientific support for researchers’ personal beliefs.
The problem that quantitative results in published articles cannot be trusted is now widely recognized and has been called a crisis of confidence a credibility crisis, or the replication crisis. However, the problem also exists at the meta-level when questionable published results are combined into a meta-analysis. Don’t get me wrong. Meta-analysis, like all statistical models, are not wrong. They are only wrong when incompetent researchers use these tools without understand how they work and what assumptions these models make.
Meta-analysis is easy to understand and perform when all data are available. We simply combine summary statistics to reduce sampling error and get a more precise estimate of the population effect size. Instead of running one study with N = 1,000 participants, we combine data from 25 studies with 40 participants. The result is practically the same. However, in psychology, the 25 published study are only a fraction of studies that were conducted and produced a significant result (Sterling et al., 1995). This means the effect sizes in the studies are inflated by publication bias and the same bias leads to an inflated effect size estimate in the meta-analysis. Thus, normal meta-analysis that ignore bias are as useful as a wet tissue paper on a 40°C (104°F) day in the middle of a parking lot at noon.
The solution to this problem is to use fancy statistical models that promise to correct for these biases and reveal the truth hidden in a pile of selected and p-hacked studies. The simple truth is that this goal is as attainable as making gold from base metals. However, as readers also do not understand these models and the problem of using them with uninformative data, the results are now routinely included in meta-analytic articles, if only as a sensitivity analysis that can be dismissed if it shows inconvenient or strange results.
Before I show how silly bias-correction of biased literature is, I need to present an example to show that I am not attacking a strawman model of bad meta-analysis. The example comes from a recent meta-analysis of studies that examined the influence of mortality salience on feelings, attitudes, and behaviors (Chen et al., 2025). The meta-analysis is notable for its attempt to deal with publication bias. The title even mentions publication bias in a clever way “Managing the Terror of Publication Bias.” The authors also shared their data. So, the only problem is that they did not consult with experts to make sense of their findings.
Figure 1 shows a simple histogram of the effect size ESTIMATES – these are estimates in small samples with enormous sampling error, not the actual effect sizes without sampling error.
Figure 1.
The most important observation for this blog post is that there are hardly any effect sizes below zero. To understand why this is important, it is important to understand the meaning of the sign of an effect size. In an original study, the sign has no meaning. For example, the height differences between people with XX and XY chromosomes can be positive or negative depending on the coding of XX as 0 or 1 and XY as 1 or 0, respectively.
For a meta-analysis, however, the sign becomes meaningful. In a competently conducted meta-analysis, the sign reflects the substantive hypothesis of a study. If mortality salience is coded as 1 against a control condition coded as 0, and the theory predicted an increase in a dependent variable, a positive sign implies that the result was consistent with the prediction. If the prediction implies a decrease in the DV, a negative sign is consistent with the theory, and the sign has to be reversed. Thus, if researchers mostly make correct predictions about the direction of an effect, we would expect mostly positive signs. Sampling error can still produce negative means in studies, even if the true effect is in the predicted direction, but how often that happens depends on the strength of the effect.
Now we are in the position to make sense of Figure 1. Only 2% of the effect size estimates. There are two possible explanations for this finding. Either TMT studies mostly produce positive results because most studies have true effects or there is selection bias and results that contradict theoretical predictions are not published (a third option would be coding mistakes, where coders code all results as positive and ignore substantive hypotheses).
What happens when these data are analyzed with bias-correction models? It depends on the model. The PET/PEESE model regresses effect sizes on the sampling error under the assumption that all studies have a common effect size and that larger samples are less biased.
The reanalysis that produced Figure 2 reproduced the published estimate of -.114 standard deviations. Thus, even though there are hardly any negative results in the data, the average study is supposed to have made a prediction in the wrong direction because that is what a negative mean means. An analysis that removed the 10% largest effect sizes, produced a positive estimate of .29 standard deviations. It is interesting that removing strong results increases the average. This shows that the results depend on assumptions about the amount of bias for different effect size estimates. Here the largest effect size estimates come also from the smallest studies (N < 10).
The point estimate of .29 should not be confused with the true effect size in each study. After removing sampling error, there is still considerable variability in the effect size estimates that can be quantified with the standard deviation, assuming a normal distribution. The estimate is tau = .40. This also makes it possible to create a prediction interval – a confidence interval for the hypothetical population effect sizes . To get a 95%CI we roughly multiply tau by 2 and get a range of values around the point estimate from .29 – ,80 to .29 + .80. Thus, any particular TMT study could have an effect size anywhere from -.51 to + 1.09. In terms of Cohen’s classification of effect sizes the effect sizes range from a moderate negative effect size to a strong positive effect size. In other words, the data are not telling us anything that we did not know before we ran the analysis. Terror Management effects may sometimes emerge as predicted, sometimes with surprising opposite effects, and sometimes have no notable effects, and we do not know which manipulation produces which effect.
The problem in the published article and many other meta-analysis is that the heterogeneity in effect sizes after taking random sampling error was ignored. The point estimate is only needed to center the prediction interval. The real information is the wide range of possible population effect sizes that are consistent with the model’s assumptions and the data. Every outcome except large negative effect sizes is possible.
Regression models have many limitations and even the developer of this approach has warned against the use of this model for highly heterogenous data (Stanley, 2017). A model that is more suitable for heterogeneous data is the weight-function selection model (Vevea & Woods, 2005). However, this model requires assumptions about selection bias. Chen et al. fitted a model that assumes different selection bias for significant negative results and non-significant results. Importantly, their model assumed the same amount of selection bias for negative non-significant results and positive significant results. This specification is important because Figure 1 shows that there are few negative effect sizes. A better way to see the problem here is to convert effect sizes and sampling error into z-values (z = effect size / standard error) to distinguish between non-significant (z < 1.96) and significant ones.
Figure 3 shows clearly that there is selection against negative results. Sampling error alone cannot explain the drop in effect size estimates from just above zero to just below zero.
The published article reports an estimated average effect size of .36. The model also estimated that only 26% of non-significant results were included in the meta-analysis. In other words, 74% were missing due to to publication bias. Finally, the model estimated that the population effect sizes had high heterogeneity, tau = .71. This leads to a very wide prediction interval around the point estimate of .36 ranging from .36 – 2*.71 to .36 + 2 * .71, which is -1.01 to 1.73. In other words, the model does not even exclude strong negative effect sizes as possible outcomes.
However, this model is misspecified because it ignores that selection against negative non-significant results is stronger than selection against non-significant positive results. I therefore ran the model again with an additional step at p = .5 (one-sided) that separates positive from negative results.. Consistent with the pattern in Figure 3, the model shows stronger selection against negative results (weight = .01, selection 1-weight = 99%) than for nonsignificant positive results (weight = .30, selection bias 70%). This improved model, however, produced a negative estimated average effect size of -.34. It also further increased the estimate of heterogeneity to tau = .96.
To understand this behavior of the model (I am more of a model analyst than a psycho-analyst), we need to understand the model’s assumption about the distribution of the unobserved population effect sizes. The model assumes an unobserved normal distribution, but the data are a truncated distribution at zero with no meaningful negative values. The model therefore fits the positive range of a normal distribution to the observed positive values. If this distribution is very wide, the normal has a large standard deviation and the model extrapolates it into the negative range. This leads to the wide prediction interval ranging from -1.34 to 2.58. Importantly, the negative range is entirely based on distribution assumption of the model . Changing the distribution assumption would change the results.
So, we have to think about the distribution assumption. When studies are more or less identical, there may be some extra variation in population effect sizes aside from sampling error. This variance can be approximated with a normal distribution (Hedges & Vevea, 1996). But when the set of studies has effect sizes ranging from 0 to 2, this is no longer plausible. If the average effect size is small and heterogeneity is large, a normal distribution implies that many substantive hypotheses have the wrong sign, but that is not really plausible. Many studies may have no real effect or really small ones, but it is harder to argue and to believe that researchers often get the sign of an effect wrong, especially when there is a real effect. Reminding people of their death makes them afraid is a reasonable hypothesis, and it would be surprising if studies show the opposite result.
In short, the weight-function model is not wrong, but applying a model that assumes a normal distribution to highly heterogeneous data is wrong. The model predicts many negative results that do not match any observed results. It could be selection bias, but it could also be a false distribution assumption. What to do?
A reasonable approach to make sense of results from the selection model is to focus on the positive side of the distribution. With normal distributions it is easy to get other statistics like the mean of only positive results (or any other subset of studies). We can therefore ignore studies with false substantive hypotheses and focus on studies where researchers made correct predictions about the sign of an effect (H1 is true).
With a mean of -.34 and tau = .956, we get a conditional mean for studies in which H1 is true of .65 standard deviations (a medium to large effect size) with tau of .52. As the lower bound is zero, we only need the upper bound and get .65 + 1.96 * .52 which is 1.67. This would suggest that many studies have strong effect sizes, which seems to contradict the estimated center of the distribution at -.34. This shows how meaningless these point estimates are when heterogeneity is large.
Unfortunately for terror management researchers the truncated moments are not going to rescue their literature because they are hypothetical. The reason is that we are conditioning on an unknown parameter, namely the condition that the hypothesis was true, but for any particular study we do not know whether H1 is true or not. So, the correct way to formulate this result is “if you can identify a study design in which terror management theory makes the right prediction and you can get a fairly precise estimate of the true effect size, you can expect a moderate effect size estimate.
What the weight-function model does not provide is a bias-corrected estimate of the positive effect size estimates in the dataset. The mean of the full distribution includes negative results that were either removed or never obtained. The truncated moment estimate conditions on the unknown status of the null-hypothesis. One is likely too low and the other is likely to high, but neither is conceptually the estimate we want. The average population effect size positive studies that corrects for the selection of nonsignificant results.
To summarize, state of the art meta-analyses in psychology try to deal with the terror of publication bias, but fail to do so. The main reason is that the statistical models that are available do not match the data. They were designed for meta-analysis of close replications with small variation in true effect sizes. They were not intended to be used for meta-analyses of diverse paradigms with large heterogeneity. Other methods that were developed after the replication crisis like p-curve and p-uniform have the same limitation. They work when heterogeneity is small, but they do not work for meta-analyses of diverse studies that are only loosely related by a common hypothesis.
Z-Curve to the Rescue
The quote “Insanity is trying the same thing and expecting a different result” has been attributed to Einstein. Even if that attribution is false, the insight is right. The problem with meta-analytic models is that they try to estimate a single number. This makes sense when the goal is estimation of a single population effect size, but not when every study has a different population effect size.
When we have a heterogenous literature, we need to face heterogeneity head on, and not hide it in some test that is reported and ignored. There is also heterogeneity around the estimate, p < .05. We need to see how much heterogeneity there. But to do that, we first need a model that can deal with heterogeneity without making unrealistic assumptions about the distribution of population effect sizes.
With a fresh look at the problem, we can look to other research areas that have addressed the problem of heterogeneity in effect sizes and large uncertainty about effect sizes of a specific result. Genomics tests millions of DNA segments (SNPs) and tries to find a few segments that show promising results. The goal here is to find the needles in the hey stack rather than averaging across millions of segments that have no relationship with a phenotype. As selection for the strongest observed effects leads to inflated estimates, models are needed to correct for this inflation. However, these corrected estimates are still tight to actual observed results rather than claims about some unobserved distribution of effect sizes. That makes it possible to identify specific segments in the observed data with promising results.
The same logic can be applied to meta-analysis. The goal is no longer to make claims like “the average population effect size is zero” or “the range of plausible effect sizes ranges from -1 to 1.” the goal is now to say “these studies show convincing evidence with meaningful effect sizes.”
One statistical model that can be used to answer this question is zurve (Brunner & Schimmack, 2020; Bartos & Schimmack, 2022). With a few modifications, z-curve can be used for directional meta-analysis where the sign of an effect matters. Rather than fitting z-curve to absolute z-values that ignore the sign and using folded normal components, z-curve can use truncated z-values and truncated normal components. When the model is fitted to only significant results, the difference is minor. More importantly, z-curve estimates of power can also be used to compute bias-corrected effect sizes (Efron, 2005). The reason is that power is a function of effect size and sampling error, so we can use the inverse normal to convert power into a corrected z-value and then multiply it with the sampling error to get a bias corrected effect size. The main challenge is to estimate the sampling error for unobserved non-significant results because their sample sizes are unknown. A simple approach is to use the sampling errors of the just significant results as an approximation. A weighted average of these estimates is the estimate of the true average effect size for the population of studies with positive results before selection for significance. Negative results that are observed are discarded.
Figure 4 shows the results of a simulation study in which the true average power before selection is known. The simulation modeled a beta distribution and graded selection bias. This is important because the weight-function model does well when its assumptions are met. The problem is that the assumptions are are untestable and often questionable. For example, we can simulate a literature with a mean of zero and tau of .4, but this simulation implies that a theories predictions are no better than a coin flip. Once researchers make better predictions, the normal assumption no longer holds.
While z-curve estimates are not perfect, they are conceptually meaningful and closer to the truth than either of the weight-function model’s estimates. We can now apply the model to the TMT data, excluding the few (2%) negative estimates.
The z-curve shows clear evidence of selection bias (the red dotted line is above the light purple bars of the nonsignificant results. However, the EDR estimate of 35% suggests that studies have on average 33% power to produce a significant result. Moreover, an EDR of 33% implies that no more than 10% of the significant results can be false positive results. Even the lower limit of the EDR confidence interval, 20%, allows for only 20% false positive results. This would suggest that many studies, especially significant ones, produced evidence for a true hypotheses with an effect size in the right direction. We can now also quantify the typical effect size. The overall effect size estimate is .63, 95%CI [.46 to .68]. Moreover, we can quantify the average for different ranges of z-values. The average increases from .40 for z-values between 0 and 0.5 to effect sizes greater than 1 for z-values greater than 4.
This finding is surprising, to say the least, because typical effect sizes in psychology are around d = .4 and rarely greater than 1. Before TMT researchers start celebrating, we have to reconcile these findings with the z-curve analysis published in the TMT article.
The z-curve looks notably different in that it does not have a long tail of high z-values. As a result, the EDR estimate is much lower, .08, and the 95% confidence interval includes alpha, 5% to 17%. This implies that there is no t enough evidence to reject the null-hypothesis that all significant results were obtained without a real effect, average effect size: zero, even for z-values greater than 4. So what is it? Is the average effect size close to zero or greater than 1?
To understand the different results, it is important to know that the published z-curve used a different coding of studies than the effect size meta-analysis. I fitted z-curve to these z-values and computed effect size and sampling error estimates from the z-values and degrees of freedom, assuming between-subject designs with equal cell sizes.
The plot is scaled to show the full distribution in the range of non-significant results. The model estimates reproduce the published results. The EDR is 7%, 95%CI = [5%, 17%]. The plot also shows local power for z-values from 0 to 3 stays low. Studies with z-values greater than 4 have acceptable local power but contrary to the previous z-curve, there are hardly any studies. This published z-curve produces dramatically different average effect size estimate, .12 95%CI = .02 to .19. The results also imply much lower heterogeneity because there are hardly any studies with strong evidence (z > 4) and large effect sizes.
Applying the weight-function selection model produces roughly the same results. The average effect size estimate is d = .17, and heterogeneity is small, tau = .17. Now the PET regression result also agrees, intercept = .05, tau = .19.
In conclusion, careful examination of this meta-analysis shows several problems. First, the data were coded inconsistently and different models were given different data. As it turns out, the effect size coding was wrong because F-values were coded as t-values, which dramatically inflates effect size estimates. Second, inconsistent results focused on the point estimate of models, but the point estimate is irrelevant when data are highly heterogenous (due to coding mistakes). Properly interpreted, all models suggest high heterogeneity that allows for large effect sizes among positive results. However, when the data are properly coded, the results show weak evidence that any study produced real effects, a high false positive risk, a small average effect size and small heterogeneity. These results change the final conclusion in the article.
“Given the conflicting findings that emerged across tools and the inherent trade-offs associated with each tool, we caution researchers against drawing firm conclusions about the evidential value of literature through any single analytic tool.”
Correction: The results are consistent and show that most studies provide no evidence for an effect because most effect sizes are small and studies had low power to detect or estimate these effects.
“PET-PEESE can underestimate the effect size when there is publication bias and when p-hacking is present (Carter et al., 2019), which are two conditions likely affecting the literature.”
Here bad research practices are used as an excuse to dismiss the most negative result without mentioning the real problem There is no “effect size.: there is only an average effect size and regression models still allow estimation of heterogeneity that was large in the data the authors used. Even a negative average can be consistent with many true positive effects when heterogeneity is large.
Z-curve can be a powerful tool for inferring the overall composite z-score distribution of a heterogeneous literature. However, unlike the other analyses included in this study, z-curve has not been as thoroughly evaluated by independent researchersso the statistical properties for its power estimates remain under explored. Furthermore, its power estimates are subject to the usual theoretical objections to estimating power from a fixed sample of data (for a recent commentary, see Pek et al., 2022).
This statement ignores that z-curve has been thoroughly evaluated by extensive simulations studies that have been reproduced by the editorial team during an open peer review process. The same cannot be said about the other methods that have not been vetted as rigorously or failed to do well in some conditions (Carter et al., 2019). The reference to Pek is also misleading which has been addressed in several rebuttals to this unfounded claim (Schimmack & Soto, 2026; Soto & Schimmack,2026) with no rejoinder by Pek.
The higher conditional power estimate therefore suggests some evidential value in published studies that yielded significant findings.
The authors are referring to the ERR estimate of 22% [16% , 37%]. Suddenly Pek’s criticism of z-curve is no longer relevant. More importantly, this finding implies that an exact replication of a study with a significant result has a 22% chance of a successful replication outcome. This is abysmal and one of the lowest ever found, not a cause for optimism. Surely reminders of mortality will sometimes have an effect on something, but a research program that uses different designs with an average power of 22% will not be able to identify when a manipulation works or when it is just a chance finding. In fact, the upper limit of the DR estimate is 100%. Thus, these weak studies fail to reject the hypothesis that all studies are pure noise.
“The selection models provide evidence for a small effect consistent with the MS hypothesis… We suggest that the average effect of the literature may be within the range estimated by the selection models and WAAP-WLS (i.e., r is around .18), although this average may have resulted from a mix of effects, many of which are higher than .18, and many of which are lower than .18.
The average estimate is too high once we correct for the coding mistakes. The real effect size is half of this (d = r / 2), and heterogeneity is small. This is the most conesquences conclusoin. Rather than having evidence of a wide range of positive effect sizes, we have evidence that most effect sizes are small and too small to study with the typical sample sizes of this literature.
We encourage future preregistered replications of the MS hypothesis to use smaller es timates of effect size (i.e., r = .18).
This inflated effect size estimate will only lead to a replication failure. Given the weak evidence in this literature, it may be better to start a new credible research program about coping with awareness of one’s own mortality than to invest more resources into this failed paradigm with questionable manipulations and dependent variables.
Though on their face the liberal and conservative interpretations feel contradictory, some observations are uncontroversial. The first observation is that the TMT literature consists of highly heterogenous.
Even this conclusion turns out to be false when the proper data are analyzed. Heterogeneity was caused by coding mistakes and practically vanishes when the correct data coded by the authors were analyzed. The authors did not notice that their data were inconsistent, even though a simple comparison of the z-values would have shown the discrepancy. It is natural for humans to make errors, but errors also reveal something about the person who committed the error. In this case, it reveals a lack of understanding of the methods, their assumptions, and why they may produce inconsistent results. Here inconsistency was attributed to properties of the models when the real source were inconsistent data. In the future, meta-analysts should not just report inconsistencies, but also try to explain them. That requires understanding of the tools that they use.
For our entire universe of studies, heterogeneity is estimated at τ = .72 under the selection models, which means that for the estimate of g = 0.36 (r = .18) for the entire literature from the selection models, 95% of the effects underlying studies of MS hypothesis, assuming a normal distribution of g, fall between g = −1.05 and 1.77 (or r = −.47 and .66); an extremely wide range of possible effect sizes arising from differences in study design.
The problem with this wide range of population effect sizes is the assumption of a normal distribution. Even if no negative results are observed, the assumption leads to the conclusion that negative effects were obtained but suppressed. But researchers are flexible and it is more likely that they would change the prediction in the direction of a significant result (Kerr, 1998). Thus, it is highly likely that the predicted negative results are phantom studies that do not exist. These predicted effect sizes surely do not correspond to the positive estimates in the dataset.
With these observations in mind, we conclude that there must be some nonzero underlying effects in the studies we examined.
That sounds more reassuring than it is. We have over 800 results and some of these are not false positives. Great, now what? We do not know which of these results are true or false positives. So, we haven’t really learned anything about mortality awareness from this meta-analysis. Fortunately, the analysis of the data without the coding mistake is more conclusive. Terror management research is an example of a pathological science. Researchers conduct studies but never learn form their data because they find a way to keep their theory alive. A proper analysis shows that we can put this literature to rest. That is ok. The history of science is filled with failures. It is also filled with examples where researchers are unable to learn from their errors. However, science moves on and experimental social psychology with little priming manipulations will be a little footnote in the history books.
P.S. And z-curve works and can now also estimate effect sizes.
To make sense of empirical data, researchers need statistical models, and the choice of model can influence conclusions. It is therefore important to be aware of the assumptions a model makes and, ideally, to test whether those assumptions hold in a particular dataset. This simple truth applies even to the choice of how to describe the average of a variable. As most students learn in a first statistics course, the mean and the median are interchangeable for symmetric distributions but not for skewed ones. In that example, the assumption is easy to check: test the symmetry of the data and pick the better summary.
With more complex models, testing assumptions is harder, and some assumptions cannot be tested at all. In these contexts it is common to leave the influence of assumptions unexamined. It is also awkward to phrase every conclusion as a conditional statement — “assuming the model’s assumptions hold, …” In psychology, correlational researchers are routinely required to flag that any causal claim is conditional on assumptions, or else they are told not to draw conclusions at all. Model-based results, by contrast, are often presented without a careful discussion of the model’s assumptions.
An increasingly popular model for meta-analyses is the step-function selection model (Vevea & Woods, 2005). Like other selection models, it assumes that nonsignificant results often remain unpublished while significant results are published. Other biases may also influence which significant results appear, but most selection models make the simplifying assumption that selection for significance is the key driver of publication bias.
Where selection models differ is in their treatment of nonsignificant results. Models like p-curve and z-curve fit only the significant results and do not condition on the observed nonsignificant ones; they avoid making any assumption about how selection shaped the nonsignificant results. Step-function models instead use the nonsignificant results directly. This can yield more information, but only at the cost of an additional assumption: that the selection process takes a specific form. That extra assumption is the subject of this post — and, as with symmetry and the mean, it is one we can test.
To use non-significant results, step-function models make two assumptions that have to be true to correct for selection bias.
Non-significant and significant effect size estimates come from the same normal distribution of population effect sizes.
Selection bias within a step (a range of p-values) is independent of the p-value.
Taken together, these two assumptions make predictions about the observed p-values within a step. This prediction is easier to test when the p-values are converted into z-values. Essentially, selection bias should influence the frequency of results in a step, but not the shape that one would see without selection bias. This prediction has never been tested, but it is possible to do so using a z-curve plot and z-curve analyses of the data.
Here I use actual data from a meta-analysis of Terror Management Studies that used the step-function model (Chen et al., 2025).
A simple histogram of the z-values shows that there are hardly any negative results (Figure 1). However, Chen et al. specified a model that assumes equal selection bias for all z-values between -1.96 and +1.96. Accordingly, the data imply that there really were more non-significant results with a positive effect than a negative effect. This alone implies that the mean of the population effect sizes is positive. Indeed, the model estimated a mean standardized effect size of .347 standard deviations.
An alternative possibility is that there is a stronger selection bias against negative non-significant results than for positive non-significant results. This assumption can be tested by splitting the step at z = 0 (p = .5, one-tailed) and estimating selection bias separately for positive and negative results. This model produces a dramatically different estimate of the average effect size, suggesting that studies, on average, more often produced a negative result (contrary to predictions) than a positive result, standardized mean difference = -.336.
The change in the sign of the estimate shows how sensitive estimates are to model assumptions. I used z-curve (Brunner & Schimmack, 2020; Bartos & Schimmack, 2022) to test the assumption that the distribution of the positive non-significant results is consistent with the assumptions of the step-function model.
Z-curve fits a finite-mixture model to the distribution of the significant z-values and predicts the distribution of non-significant results from the weights of the mixture components (Figure 2). The z-curve plot shows clear evidence of selection bias. That is there are more observed significant results (66%) than the model predicts (31%). However, the important question is whether the distribution of the non-significant positive z-values matches the predicted distribution shape (the red dotted line in Figure 1).
Visual inspection suggests that this is not the case. There are more observed non-significant results close to the significance criterion than close to zero. In contrast, the predicted pattern shows a flat and then decreasing shape from 0 to 1.65 (values between 1.65 and 1.96 are inconclusive because they are often reported as marginally significant results).
To complement visual inspection with a statistical test, z-curve (version 3.85) compares the distributions while holding the overall density constant (i.e., removing selection bias). A negative difference indicates that there are more z-values close to zero than predicted. A positive difference indicates that there are more z-values close to the significance criterion. A bootstrap confidence interval is used to test for statistical significance.
The median difference was z = .15, 95%CI [ .08, .38]. The mean difference was z = .10, 95%CI = [.05 to .26]. This finding is consistent with the observation that observed z-values close to zero are relatively rare. In short, the observed distribution of the non-significant results is inconsistent with at least one of the assumptions of the step-function model.
The test does not reveal which assumption is violated, and in these data both may be. The depletion of nonsignificant results near zero suggests that selection bias varies within the step, contrary to the assumption that selection is constant within a p-value range. Separately, the near-absence of negative results is hard to reconcile with the model’s estimate of large heterogeneity (tau = .96), which produces a prediction interval for population effect sizes from d = −2.20 to 1.54. A lower bound of −2.20 implies that a substantial share of studies have true effects opposite to the prediction — that mortality primes often produce the reverse of the expected effect. That is implausible; more likely, the negative tail is an artifact of assuming a normal distribution. The shape test cannot identify the true distribution of population effect sizes, but it forces meta-analysts to confront the distribution assumption rather than leave it implicit.
To demonstrate that the shape test works, I used the model weights of the actual data to simulate data without publication bias. I then removed 70% of the non-significant results without changing the distribution of the non-significant results (Figure 2).
Visual inspection shows that the distribution of the non-significant results is more similar to the predicted distribution. The shape test results are consistent with this observation. The median difference is -.05, 95%CI [-.10, .18]. The mean difference is -.04, 95%CI [-.07, .12]. Both confidence intervals include zero.
In sum, step-function models make assumptions about the distribution of population effect sizes and the selection of non-significant results to use observed non-significant results in the estimation of the amount of bias and the average and standard deviation of the population effect sizes. The advantage is that more data reduce random sampling error. The disadvantage is that violations of assumptions can introduce systematic biases. This disadvantage has been neglected in applications of step-function selection models. The ability to test the assumptions has several benefits. First, it draws more attention to the fact that assumptions influence model estimates. Second, shape tests help readers to evaluate the plausibility of the model and to question estimates that are based on data that are inconsistent with assumptions. Third, assumption tests can resolve inconsistencies between estimates obtained with different models. Results based on models that make fewer assumptions or pass assumption tests are more credible than those based on models that do not meet assumptions.
In conclusion, meta-analysts have many tools to analyze their data that often produce inconsistent results. To make sense of these inconsistencies, it is important to understand why they produce inconsistent results. Violated assumptions are one plausible reason and undermine the plausibility of results of these models.
Meta-psychology was born from a simple observation: the way psychologists used significance testing created a distorted literature. Researchers treated p < .05 as a license to publish and p > .05 as a reason to abandon a finding. Journals rewarded significant results, reviewers demanded them, and authors learned to find them. The result was predictable: literatures stuffed with too many significant results, exaggerated effect sizes, and too few honest failures (Sterling, 1959; Sterling et al., 1995).
This critique is now familiar. Null-hypothesis significance testing, reduced to a dichotomous decision rule, encourages bad scientific behavior. It turns evidence into a yes-or-no ritual, treats p = .049 as a discovery and p = .051 as a non-event, and rewards selective reporting. Meta-psychologists have made this point repeatedly, and largely correctly.
But there is an irony. Having criticized psychologists for using a dichotomous significance test to decide which original findings count, meta-psychologists often reach for the same logic to decide whether a literature is biased.
The original sin was this: p < .05 means the effect is real.
The meta-analytic version becomes this: p < .05 means publication bias is present.
The form of the reasoning has not changed. Only the target has moved up one level.
Why this fails is clearest if we ask what a significance test can ever legitimately buy us. The most charitable defense of significance testing is that a significant result may carry information about the sign of an effect: it can tell us which direction is more plausible, even when it says little about magnitude (Jones & Tukey, 2000). That defense collapses for publication bias, because the sign is known before we collect a single study. Selective reporting favors significant results; it does not run the other way. A test whose only defensible output is a direction we already know contributes little.
What it contributes instead is a verdict that is uninformative in both directions. A significant bias test conflates magnitude with detectability: in a large literature, a trivial and harmless amount of selection can reject the null. A nonsignificant bias test conflates small bias with low power: in a small literature, severe selection can easily fail to reach significance. Either way, the binary outcome tells us little about the quantity we care about. “There is publication bias, p < .05″ is, to borrow Cohen’s (1994) famous example, about as useful as “the earth is round, p < .05.” And “There was no evidence of publication bias, p > .05” is akin to “The earth is flat, p > .05.”
The deeper irony is that meta-psychologists have relocated the mistake they diagnosed. Original researchers treated significance as a discovery machine. Bias researchers sometimes treat significance as a bias-detection machine (Siegel et al., 2021). The error is identical: a difficult inferential problem is compressed into a binary decision.
Some have pushed the argument further, claiming that tests for publication bias are useless (e.g., Simonsohn, 2014). But the folly of nil-hypothesis testing, which incidentally undermines p-curve as much as many other significance-based methods, is not a reason to ignore publication bias. We do not abandon original research because the significance ritual is empty (Cohen, 1994). We replace the ritual with something more informative.
The reform for original research was to report effect-size estimates with confidence intervals that express uncertainty. The reform for bias detection should be the same. The goal is not to decide whether bias exists, but to estimate how much is present, how uncertain that estimate is, and whether the amount of bias consistent with the data changes the substantive conclusion.
Some publication-bias methods already estimate quantities of this kind, or carry the information needed to. Yet in practice that information is discarded, and the result is reduced to whether a test was significant or a method “detected bias” (Siegel et al., 2021). And no common metric for the amount of bias has been widely adopted.
The most natural metric is the excess of significant results. If a literature reports significant findings 80% of the time but the true probability of producing significant results is between 20% and 40%, we have clear evidence of substantial bias.
This is why the amount matters more than its presence. Bias can be easy to detect yet too small to change any conclusion, or large enough to overturn a conclusion yet impossible to detect in a small set of studies, where these tests have the least power (Renkewitz & Keiner, 2019). A binary test cannot tell these cases apart; an estimate with an interval can.
In short, meta-analysis needs the same methodological reform that original research needed. It is time to abandon the nil-hypothesis ritual and replace it with estimation: estimate the amount of publication bias, quantify the uncertainty with confidence intervals, and evaluate whether conclusions remain credible after adjusting for the plausible levels of selection.
Fortunately, unlike unpublished primary studies hidden in file drawers, the data behind published meta-analyses are often available or recoverable. That makes it possible to reexamine decades of meta-analytic conclusions and ask the question that matters: not whether publication bias can be detected, but whether the amount of bias compatible with the data changes what we should believe.
The past decade has not been kind to experimental social psychology. Study after study failed to replicate and entire literatures have turned out to be built on nothing (a.k.a. statistical noise mining).
“Another day, another idol falls. This one has been teetering for years, so the collapse didn’t come as a shock. But that doesn’t make it any less painful.” (Michael Inzlicht).
It all started with a leading journal publishing an article with the crazy claim that people can foresee the future and practicing after a test can improve exam scores (Bem, 2011). This claim was quickly revealed to be false (and possibly a hoax, Gelman) after a big replication study failed to show the same results (Galak, J., LeBoeuf, R. A., Nelson, L. D., & Simmons, J. P., 2012).
In a media interview Bem explained that his experiments were never meant to be taken seriously. (Daniel Engber, 2017, Slate).
“If you looked at all my past experiments, they were always rhetorical devices. I gathered data to show how my point would be made. I used data as a point of persuasion, and I never really worried about, ‘Will this replicate or will this not?’
While the past decade has not been good for experimental social psychologists, it has produced a new group of psychologists to examine the causes of the replication crisis in experimental social psychology. As they look at the practices of research psychologists, they are meta-psychologists, psychologists who study other psychologists.
One of them is Blake McShane, who did his dissertation on statistical models to analyze time-series data (McShane, 2010). Given his background in statistics, managerial science, applied economics, and marketing, it is fair to say that he entered this field without first-hand experience of research practices that produced the replication crisis. He also does not cite seminal papers that foreshadowed the crisis by Cohen (1962, 1990, 1994). Instead, his main approach to examining meta-psychological questions appears to rely on his expertise in conducting simulation studies (McShane & Böckenholt, 2014, McShane, Böckenholt, & Hansen, 2016, 2020).
The problem with these simulation studies is that they repeat the same problems that plagued experimental social psychology at the meta-level. Just like Bem’s studies are not empirical tests, but rhetorical devices, McShane’s simulations are rhetorical devices to illustrate a point that does not require simulation evidence, namely.
[models] perform reasonably well in the setting for which they were designed, …[but] they are sensitive to deviations from their model assumptions.
In the 2016 article, the simulations violated assumptions of models that assume homogeneity and they failed. However, the simulations met the assumptions of another model and (no surprise) it worked well. However, McShane did not cite an earlier study that showed the model also has problems when its assumptions are not met (Hedges & Vevea, 1996).
Later simulation studies further confirmed that McShane’s preferred model does not work so well under realistic conditions (Carter et al., 2019), a finding not cited by McShane et al. in 2020. Pressed on this point that his simulations favored his preferred model, he might reply
“If you looked at all my past simulations, they were always rhetorical devices. I created conditions to show how things work when assumptions are met. I used simulations as a point of persuasion, and I never really worried about, ‘Does this apply to real data’ ”
In conclusion, a simulation that shows a model works when its assumptions are true and does not work when its assumptions are false is merely a demonstration, not an evaluation of a model under realistic conditions.
Cookie Consent
We use cookies to improve your experience on our site. By using our site, you consent to cookies.