
Brief Introduction
The replication crisis has not been kind to priming research. Social priming has become the poster child for the kind of replication failures that are expected when researchers conduct underpowered studies, use flexible analyses, and selectively publish results that crossed, or nearly crossed, the conventional threshold for statistical significance, p < .05. The problem is not mysterious. If many small studies are run and only the successful ones enter the published literature, the published record will exaggerate effects, hide failures, and create the illusion of a robust phenomenon.
Yet defenders of social priming continue to insist that priming is real and robust. Their main defense is no longer a stream of new, well-powered, preregistered experiments showing that classic priming paradigms work. Instead, the defense has shifted to meta-analysis. The argument is that even if individual studies are weak, the combined literature still shows an effect. Dai et al. (2023) make this argument explicitly. They conclude that the field has moved beyond the question of whether priming exists and should now focus on mechanisms.
That conclusion is not supported by the evidence. A meta-analysis of p-hacked studies is not a cure for p-hacking. If the primary studies were produced by low power, selective reporting, flexible analysis, and publication bias, then the meta-analysis inherits those problems. A significant average effect does not show that priming is robust. It does not identify a reliable paradigm. It does not show that effects occur outside awareness. It does not explain why preregistered and direct replication attempts have often failed. And it does not justify moving from existence questions to mechanism questions.
An empirical science needs paradigms that can reliably produce the phenomenon under specified conditions. If behavioral priming is real and robust, proponents should be able to demonstrate it prospectively in well-powered studies. Cohen made this point long ago: science requires studies with a decent chance of detecting the effects they are designed to test. If the true effect is around d = .30 or d = .40, then the classic small priming studies were severely underpowered. Underpowered studies can generate significant results, but those results will be selected, inflated, and difficult to replicate.
This raises the obvious question: if priming works, why are its strongest defenders not still doing programmatic priming research? A robust phenomenon should generate active research programs: stable paradigms, preregistered replications, planned moderator tests, and cumulative theory development. Instead, much of the contemporary defense of priming relies on retrospective meta-analyses of old, heterogeneous, mostly nonpreregistered studies. Actions speak louder than words. If priming is robust, show that it works in the lab.
I have tried to publish a critique of Dai et al.’s meta-analysis in traditional psychology journals. Psychological Bulletin, which published the original meta-analysis, was not interested in publishing a correction of its conclusions. PSPR was also not interested in a paper that challenges the integrity of the evidentiary practices behind one of social psychology’s most famous literatures. Fortunately, journal gatekeeping does not prevent readers from evaluating the arguments for themselves.
The following list summarizes the main problems in Dai et al.’s meta-analysis and explains why their conclusion that behavioral priming is a robust phenomenon is not supported by their own evidence.
Scientific Problems in Dai et al. (2023) Meta-Analysis
- A significant average effect is not evidence of a robust phenomenon.
Dai et al.’s main empirical result is a positive mean effect, d = 0.37, across 351 studies, 224 reports, and 862 effect sizes. They interpret this as evidence of a moderate and robust priming effect. But a positive grand mean only shows that the average coded effect is greater than zero. It does not show that any specific priming paradigm works reliably, that the effect is predictable, or that the literature has identified the conditions under which the effect occurs. This is the central inferential error. A heterogeneous collection of effects can have a positive mean even if many individual paradigms are unreliable, null, or false positive.
- The literature is highly heterogeneous.
Dai et al. report substantial heterogeneity in the overall analysis and therefore use random-effects models. Their results show “substantial variability across studies and contexts,” but they still conclude that the effect is robust across procedures. That is too strong. High heterogeneity means the average effect is not a good description of individual studies or future studies. Robustness would require a narrow range of expected effects, or at least identifiable and replicated boundary conditions. Heterogeneity alone is not evidence for a robust phenomenon.
- They do not report prediction intervals.
A random-effects meta-analysis with large heterogeneity needs prediction intervals. A confidence interval around the mean asks whether the average effect is different from zero. A prediction interval asks what effect size should be expected in a new study. That is the relevant quantity for “robustness.” If the prediction interval includes zero or negative effects, the literature does not support the claim that priming reliably produces positive behavioral effects. Your reanalysis indicates that Dai et al.’s estimates imply very wide prediction intervals, including negative effects, which directly contradicts the robustness claim.
- Publication bias is present by their own analyses.
Dai et al. do not find a clean literature. They report evidence of small-study bias and funnel asymmetry. Their PET–PEESE analyses indicated that small studies reported larger effects, and the PEESE-adjusted estimate was reduced to d = 0.291, although still statistically significant. A reduced but still significant average does not eliminate the publication-bias problem. It merely says that one model still leaves a nonzero mean. It does not show that the published literature is robust, nor that individual paradigms are replicable.
- Some bias-correction methods give much smaller estimates.
Dai et al. use multiple publication-bias adjustments, including PET–PEESE, Vevea and Woods selection models, and Mathur and VanderWeele sensitivity analyses. They acknowledge that selection models should be used to explore a range of estimates rather than to obtain one preferred estimate. The conclusion nevertheless emphasizes robustness. That is selective interpretation. A fair conclusion would stress the uncertainty and model dependence of the bias-corrected estimates, not reassure the field.
- Their selection models do not handle dependence among effect sizes.
Dai et al. explicitly note that Vevea and Woods selection methods do not account for statistical dependence among effect sizes. This matters because their dataset contains multiple effect sizes from the same studies, samples, and reports. If selection-bias corrections are applied in a way that ignores dependence, the adjusted estimates and uncertainty intervals can be misleading. This weakens the evidentiary basis for strong claims about robustness.
- The preregistered studies are essentially null.
This is one of the most damaging findings. Dai et al. report that only a small number of reports were preregistered, but those preregistered studies did not show the same effect pattern as the nonpreregistered literature. In your summary, the key contrast is preregistered studies around d = .02 versus nonpreregistered studies around d = .37. That contrast is fatal to a simple robustness claim. If priming were robust, preregistered studies should not collapse toward zero.
- They dismiss preregistered failures asymmetrically.
Dai et al. explain the preregistered/null pattern by pointing to the small number of preregistered studies, replication difficulties, or possible differences between original studies and replications. That may be partly true, but it creates an asymmetric evidentiary standard. Positive nonpreregistered findings are treated as evidence for priming; preregistered failures are treated as inconclusive. This is precisely the kind of post hoc protection that prevents a theory from being falsified.
- Most included studies were not preregistered.
Dai et al.’s descriptive table reports that among 224 included reports, only three reported preregistration. That means the positive average is overwhelmingly based on the older, flexible, nonpreregistered literature. This is exactly the kind of literature in which selective reporting, optional stopping, outcome switching, participant exclusions, and analytic flexibility can inflate effects.
- Participant exclusions and covariate use are common.
Dai et al. report that many reports included covariates or excluded participants. Their study-quality section treats preregistration, covariates, and exclusions as indicators relevant to possible p-hacking. These features do not prove bias in any particular study, but they are red flags in a literature where the main evidentiary claim depends on many small, flexible, nonpreregistered experiments.
- Their “quality assessment” is weak.
Dai et al. count random assignment, participant blindness, and signs of p-hacking. But random assignment was an inclusion criterion, so it cannot distinguish stronger from weaker studies. For blindness, only 37% of studies used funneled debriefing; for the rest, they state that it was hard to determine whether participants were blind, but that awareness was probably unlikely because priming effects are subtle. That is not a strong quality assessment. It assumes what needs to be shown.
- Awareness is inadequately handled.
The controversial claim is not that cues can influence behavior when participants notice them. The controversial claim is influence without awareness of the prime’s relevance or influence. Yet Dai et al.’s own descriptive table shows that only 37.82% of effect sizes came from studies with funneled debriefing, and 84.22% used supraliminal primes. This means their broad meta-analysis does not primarily test the strongest implicit-priming claim.
- Supraliminal cue effects are conflated with implicit priming.
Most studies in Dai et al. involve visible primes. Visible primes may influence behavior through ordinary cueing, demand, construal, motivation, or conscious goal activation. That is not the same as showing that behavior is guided outside awareness. Their conclusion about priming as a long-debated phenomenon slides between the broad, relatively uncontroversial claim that incidental cues can influence behavior and the stronger, controversial claim that behavior is automatically influenced outside awareness.
- The dataset combines many theoretically distinct phenomena.
Dai et al. include primes involving achievement, common behaviors, money/marketing/finance, morality/God/prosociality, motivation, sex/gender/romantic behavior, and stereotypes. They also include verbal and visual primes, subliminal and supraliminal primes, performance outcomes, donations, choices, consumption, reaction times, and other behavioral measures. A positive average across this broad mixture does not validate a single phenomenon called behavioral priming. It may simply show that many different manipulations sometimes influence many different behaviors.
- “Behavioral” versus “nonbehavioral” priming is too broad to identify mechanism.
Their core theoretical contrast is behavioral versus nonbehavioral primes. But they find no meaningful difference between the two: nonbehavioral primes d = 0.394 and behavioral primes d = 0.344, with B = 0.05, p = .14. They interpret this as suggesting overlapping processes. But a null difference between broad categories does not identify a mechanism. It may simply show that the categories are too crude, heterogeneous, or underpowered for mechanistic inference.
- Moderator analyses are weak because the relevant moderator conditions are rare.
Dai et al. acknowledge that most included studies did not manipulate the key theoretical moderators: 89% did not manipulate goal value, 91% did not manipulate goal expectancy, and 82% did not include a filler task between priming and behavior. This makes strong moderator conclusions inappropriate. A literature cannot establish mechanisms if the relevant moderators are mostly absent.
- Uneven moderator distributions reduce power.
Dai et al. explicitly state that the distribution of key moderators was highly uneven and that RVE has low power for moderators with uneven distributions. This is an important limitation. Moderator claims require interaction evidence, and interactions require much larger samples than main effects. A sparse and uneven moderator structure cannot support the claim that the field has moved from existence to mechanisms.
- Simple effects are not evidence of moderation.
Even if some simple effects appear strong in certain moderator cells, the relevant evidence is the interaction. A significant effect in one condition and a nonsignificant effect in another is not evidence that the conditions differ. If the interaction evidence is weak, underpowered, or selected, moderator claims are unstable. This is especially important because the literature is already biased toward significant results.
- Heterogeneity is treated as theoretical richness rather than an evidentiary problem.
Dai et al. repeatedly frame variability as consistent with different mechanisms or contextual dependence. But unexplained heterogeneity is not evidence for moderators. It is evidence that the effect is not yet predictable. The field cannot move to mechanism questions until it has reliable paradigms that produce the effect under specified conditions.
- Their conclusion overstates what meta-analysis can show.
Dai et al.’s conclusion says their synthesis provides “solid evidence” that priming is real, not merely a bubble of publication bias, and that it should “reassure the field.” That is stronger than the data warrant. Meta-analysis of a selected, heterogeneous, mostly nonpreregistered literature can suggest that some effects may be real. It cannot establish that behavioral priming is robust as a general phenomenon.
- The power implications are not drawn.
If the true effect is d = .29 to .37, typical small priming studies are underpowered. Many classic experiments with n = 20–40 per cell would have low power to detect effects of this size. Dai et al. do not make this implication central. If most original studies were underpowered, then significant findings are expected to be selected and inflated, and moderator analyses are even less credible.
- The positive average is compatible with many false positives.
A significant average effect does not tell us how many individual significant findings are false positives. In a heterogeneous literature with low power and selection for significance, many significant findings can be false positives even if the mean effect is positive. This is why z-curve and related evidential-value analyses are needed. Dai et al.’s focus on mean effect size does not answer the false-discovery question.
- Their methods estimate effects, not replicability.
Dai et al. assess mean effects and publication-bias-adjusted mean effects. But robustness is fundamentally about whether findings replicate. A field can have a nonzero average effect and still have poor replication prospects if studies are underpowered, selected, and heterogeneous. Their analysis does not directly estimate expected replication rates or discovery rates.
- Failed replications are not integrated as severe tests.
Dai et al. cite failed replications as part of the motivation for the meta-analysis, but the conclusion effectively downgrades their importance once the average effect remains positive. That is not a balanced evidentiary treatment. Large, preregistered, or direct replication failures should receive greater weight as severe tests of the robustness claim, not be absorbed into a heterogeneous average and then explained away.
- “Robust to publication bias” is too weak a criterion.
Their claim that the effect remains statistically significant after some bias adjustments does not establish robustness. A robust phenomenon should show reliable effects in well-powered, preregistered studies and should produce predictable effects in new studies. Surviving some bias corrections is a much lower standard.
- The meta-analysis does not identify a robust paradigm.
The key scientific question is not whether the average across 862 effects is positive. The key question is: which priming paradigm reliably produces the effect? Dai et al. do not identify such a paradigm. Without a robust paradigm, the field cannot credibly study mechanisms, moderators, or applications.
- One-off applied studies do not establish cumulative science.
The literature contains many isolated demonstrations involving different primes, outcomes, populations, and contexts. That pattern can produce a significant meta-analytic average while still lacking cumulative progress. A robust phenomenon should generate programmatic research around stable paradigms. Dai et al.’s synthesis does not show that this has happened.
- The analysis treats conceptual replication as strength, but it also weakens interpretability.
A broad conceptual-replication database increases sample size, but it also combines many different true effects. That helps reject the null hypothesis that the average is zero, but it makes it harder to know what was actually established. If the procedures differ too much, the grand mean becomes theoretically thin.
- The conclusion moves from “some effects may be real” to “the field should be reassured.”
The defensible conclusion is modest: some incidental priming effects may be real, and the average coded effect remains positive under some bias corrections. The stronger conclusion—that behavioral priming is robust and that the field can move beyond existence questions—is not supported.
- The strongest claim about implicit influence is not established.
Priming with awareness is not controversial. The important claim is that incidental cues influence later behavior outside awareness of their influence. Dai et al.’s database is mostly supraliminal, often lacks funneled debriefing, and combines many visible cue manipulations. Therefore, even if their average effect were accepted, it would not establish the stronger claim that behavior is robustly influenced outside awareness.
Conclusion
Dai et al.’s conclusion that behavioral priming is robust is stronger than their evidence permits. Their own results show substantial heterogeneity, evidence of small-study and publication bias, sparse preregistration, null effects in preregistered studies, highly uneven moderator distributions, and limited assessment of awareness. A positive average effect across a heterogeneous and mostly nonpreregistered literature may show that some incidental cue effects are real, but it does not identify a reliable priming paradigm, establish robust boundary conditions, or justify moving from existence questions to mechanism questions. Robustness requires prospective replication under specified conditions, not merely a statistically significant grand mean.
1 thought on “Priming: A Pathological Paradigm”