Category Archives: Uncategorized

Why Most Meta-Analyses (in Psychology) Could Be False: Mismodeling Heterogeneity

This blog post follows up on an introduction to meta-analyses with heterogeneous data.

Modeling Heterogeneity in Meta-Analyses – Replicability-Index

We can distinguish two types of meta-analysis. Meta-analyses of direct replication studies combine studies with the same or very similar population effect sizes. So-called fixed-effect meta-analyses produce more precise estimates of the shared population effect size, because the sampling error of the combined data is smaller than the sampling error of the individual studies.

However, psychologists often conduct conceptual replication studies that vary conditions, stimuli, and dependent variables. Meta-analyses of these studies are more challenging because the studies have varying population effect sizes (van Erp et al., 2017). As a result, the average effect size is relatively uninformative and provides no information about the effect sizes of specific studies. The standard solution has been to model this heterogeneity by assuming a normal distribution for the population effect sizes, in addition to the normal distribution of sampling error. The previous post showed that (a) this assumption can lead to biased estimates when the true distribution of population effect sizes is not normal, and (b) developments in statistics make the assumption unnecessary. Meta-analysts can therefore drop the problematic normality assumption by adopting more flexible mixture models that make no assumption about the shape of the population effect-size distribution.

What is a Positive Effect?

The main contribution of this post is to reveal a theoretical contradiction between the assumption of normally distributed population effect sizes and the way studies are coded in meta-analyses. To see the problem, it helps to ask a simple question: what is a positive effect?

The difficulty of defining a positive outcome is well recognized in medicine. A positive diagnosis usually means bad news — except when people hoping for a child learn that they are pregnant. Cochrane meta-analyses address this by coding results not only by the direction of change in an outcome (increase or decrease) but also by whether that change is beneficial or harmful: for mortality, an increase is harmful; for longevity, an increase is beneficial. For the meta-analyst, what matters is keeping the sign of an effect constant across studies, while the interpretation depends on the substantive meaning of the average. “Women live longer than men” is a positive difference if women are coded 1 and men 0, and a negative difference if the coding is reversed.

The default coding in a meta-analysis that tests a prediction is to assign signs according to the predicted direction: results supporting the hypothesis receive a positive value, and results contradicting it receive a negative sign.

This coding, combined with the assumption of a normal distribution of population effect sizes, leads to a paradoxical implication. The paradox is clearest in the extreme case where the average effect size is zero. Under the normality assumption, a mean of zero implies that half of the true effects are consistent with the prediction and half are opposite to it — that the theory predicts the direction of the effect no better than a coin flip. This implication is usually implausible, which is better read as evidence against the distributional assumption than as a verdict on the theory. Even when the average effect is small but non-zero, normally distributed heterogeneity implies that a substantial proportion of true effects run opposite to the prediction.

This state of affairs is often overlooked when meta-analytic results are interpreted. For example, Chen et al. (2025) reported large heterogeneity, with a 95% prediction interval from −1.05 to 1.77, implying that many studies have large true effects opposite to the predicted direction. Yet none of the published studies reported such results. The prediction interval implies a population of wrong-direction true effects that does not appear in the published record, and this implication went unexamined: are these effects real but suppressed by publication bias, or are they phantoms produced by an incorrect distributional assumption?

In short, coding a meta-analysis by agreement with a theoretical prediction, combined with the assumption of a normal distribution of effect sizes, predicts that many studies have true effects in the wrong direction. Because the actual data often show no such results, the prediction raises a question: do these wrong-direction effects exist and were suppressed by publication bias, or are they phantom effects manufactured by a false distributional assumption?


Modeling Only Positive Effect Sizes

When the data contain few negative effect-size estimates, the normality assumption can do more than inflate estimates of heterogeneity; it can also bias the effect-size estimate itself. The reason is that sampling error produces many negative estimates for small effects, even when the population effect is positive. If the data do not show these expected negative results, a model that assumes complete data infers that the positive results must have been produced by a larger population effect. This inflates the effect-size estimate, especially when the average effect is small.

To avoid this inflation, the truncation of the data at zero must be modeled. For example, if the true population effects follow a normal distribution centered at zero, a model that knows the data are truncated at zero expects a half-normal distribution and correctly infers that the mean of the full distribution is zero. A model that does not know the data are truncated fits a full normal distribution to the observed half-normal data and estimates a mean greater than zero.

To illustrate this, I ran a simulation with 64 conditions, varying the mean of the population effect sizes (4 levels), their standard deviation (4 levels), and the proportion of true null hypotheses (4 levels) — the same design used previously to show the problem of assuming normality when the true distribution is non-normal. The previous simulation showed that models with flexible distribution assumptions perform better than models that assume a normal distribution when the actual distribution is not normal.

Modeling Heterogeneity in Meta-Analyses – Replicability-Index

This time, the simulation selected only positive effect-size estimates. The prediction was that ashr and RMA would be biased, because they assume no selection on the sign of an effect, whereas z-curve and weightr are selection models that can be specified to model selection on sign. Weightr, however, may still be biased when the distributional assumption is violated, because its selection-weight parameters can absorb the non-normality, fitting a non-normal distribution by distorting the estimated selection process. Z-curve was expected to perform well because it makes no assumption about the shape of the effect-size distribution and can model truncated data.

The results confirmed these predictions, which follow directly from the models’ assumptions. Z-curve had the smallest error (RMSE = .046), followed by RMA and ashr (RMSE = .094), with weightr the worst (RMSE = .168).

The poor performance of the weight-function selection model is noteworthy, because it is commonly assumed to produce more credible results when data are selected. This confidence is justified only when the distributional assumption holds. When it does not, the selection model can produce worse estimates than a model that does not correct for selection at all, such as RMA. A false distributional assumption can even yield negative estimates of the average population effect when all studies have positive population effects.

Finally, it is worth noting that ashr was not designed for meta-analysis, but to analyze data from studies with many predictor variables, where the sign of an effect is often arbitrary and roughly equal numbers of positive and negative results are expected. When all of these results are available, as in genomics, the model works well — indeed, without selection it is equivalent to z-curve. The bias identified here arises only when a model built for complete data is applied to selected meta-analytic results (e.g., van Zwet et al., 2024).

In conclusion, for meta-analysis it is problematic to assume that negative effect sizes estimates are common and produced by studies with negative population effect sizes. It is therefore necessary to allow for selection based on the sign of an effect size independent of selection for statistical significance. The most suitable models for meta-analyses therefore need to assume selection for sign without assuming a particular distribution of population effect sizes. The only model that checks both boxes is z-curve.

Illustration with Actual Data

Chen et al.’s (2025) meta-analysis of terror management studies serves as a useful example, because it analyzed the data with both a random-effects model and a weight-function selection model, and because the set of studies is large (k = [FILL IN]). As noted earlier, the models that assumed normality estimated very high heterogeneity, with a prediction interval extending well into negative values — implying that a substantial proportion of studies have true effects opposite to the theoretical prediction.

The histogram shows the distribution of the observed effect-size estimates. There are many small positive estimates, but negative estimates are rare. This is not what a normal distribution would produce: a normal distribution centered on the positive mean, with the estimated heterogeneity, would place substantial mass below zero and thus predict many negative estimates. The near-absence of negative estimates therefore points to one of two conclusions — either negative results were suppressed by publication bias, or the true distribution of effects is not normal.

Chen et al. (2025) did not model selection for sign, despite the visible asymmetry around zero. I refitted the model with a selection step at p = .5 (one-tailed), distinguishing positive (p < .5) from negative (p > .5) results. This model estimated a negative average effect, d = −.336, with large heterogeneity (tau = .956). A negative mean for a literature whose observed effects are overwhelmingly positive is implausible on its face. It arises because the model can only reconcile the near-absence of negative results with a normal distribution by assuming that a large number of negative and null results were suppressed — so many that the true distribution is centered below zero, and the observed positive results are merely its selected upper tail. This is possible only if one accepts both that negative results were heavily suppressed and that the true effect distribution is normal.

The random-effects model faces a different problem. It assumes no selection and fits a normal distribution to a set of effect sizes that are visibly skewed. Fitting a symmetric normal to this right-skewed distribution inflates the mean, which RMA estimates at .81.

The ash model makes no assumption about the shape of the effect-size distribution and produces a similar mean, .80, but a smaller median, .62, reflecting the right skew visible in the observed data.

Both RMA and ash, however, take the observed data at face value, even though the data contain few negative results. Z-curve can model truncated data without assuming a normal distribution. The next figure shows the z-curve plot with local power and local effect-size estimates. This model does not assume publication bias.

Z-curve models truncation of the data at zero and makes no distributional assumption about the population effect sizes. The z-curve plot shows the histogram with the fitted model, along with local average effect sizes below the x-axis and local power estimates beneath them. Even the non-significant results are estimated to have moderate effect sizes (~.5), and the overall mean is .77.

As with ash, the median is smaller than the mean — .63 versus .77 — reflecting the same right skew. That two flexible methods independently recover both a mean near .8 and a median near .62 indicates the skew is a real feature of the effect distribution, not an artifact of either method. A random-effects model cannot represent this asymmetry, because assuming normality forces its mean and median to coincide.

Across models, then, the mean estimates agree reasonably well — z-curve (.77), ash (.80), and the random-effects model (.81) all indicate a moderate-to-large average effect — with the weight-function selection model the lone exception, producing an implausible negative estimate. For this literature, truncation and the normality assumption have a relatively small effect on the point estimate, consistent with the simulations, which showed close agreement among models when the average effect is large. This is the favorable case, however: when the average effect is small, the normality assumption’s implied negative effects and the truncation at zero have a much larger impact.

So far, all models assumed no selection for statistical significance. This is unlikely, because published articles predominantly report significant results that support theoretical predictions (Sterling et al., 1995), and Chen et al. (2025) found evidence of publication bias in this literature. Z-curve can model publication bias by fitting only the significant results and estimating the distribution of the non-significant results that were not reported.

Taking publication bias into account has two effects. First, the effect-size estimates decrease. Second, the average is now computed over a larger reference set that includes the estimated non-significant studies, which have smaller effects. The mean effect size is reduced to .52 and the median to .34. Even after this correction the distribution remains right-skewed — the mean exceeds the median — and the median of .34 is the more representative value for a typical study.

In conclusion, meta-analytic models that make different assumptions produce different results. Making sense of these differences requires testing the assumptions and preferring models that avoid assumptions that cannot be tested. In this example, the data show clear evidence of selection on the sign of the effect, further evidence of selection against non-significant positive results, and evidence that the population distribution is not normal (median < mean). A model that accommodates these features of the data — z-curve — produced lower estimates than the random-effects model, while the weight-function selection model produced an implausible negative estimate.

It is not possible to generalize from this single example to all meta-analyses, but the results raise the possibility that many published effect-size estimates are biased by false distribution assumptions, especially when publication bias is also present. To assess the extent of these biases, such literatures should be reexamined with methods that can model selection and avoid the untestable assumption of normally distributed population effects — as z-curve does.


The Irony of Testing for Bias with the Same Tools That Created It

Meta-psychology was born from a simple observation: the way psychologists used significance testing created a distorted literature. Researchers treated p < .05 as a license to publish and p > .05 as a reason to abandon a finding. Journals rewarded significant results, reviewers demanded them, and authors learned to find them. The result was predictable: literatures stuffed with too many significant results, exaggerated effect sizes, and too few honest failures (Sterling, 1959; Sterling et al., 1995).

This critique is now familiar. Null-hypothesis significance testing, reduced to a dichotomous decision rule, encourages bad scientific behavior. It turns evidence into a yes-or-no ritual, treats p = .049 as a discovery and p = .051 as a non-event, and rewards selective reporting. Meta-psychologists have made this point repeatedly, and largely correctly.

But there is an irony. Having criticized psychologists for using a dichotomous significance test to decide which original findings count, meta-psychologists often reach for the same logic to decide whether a literature is biased.

The original sin was this: p < .05 means the effect is real.

The meta-analytic version becomes this: p < .05 means publication bias is present.

The form of the reasoning has not changed. Only the target has moved up one level.

Why this fails is clearest if we ask what a significance test can ever legitimately buy us. The most charitable defense of significance testing is that a significant result may carry information about the sign of an effect: it can tell us which direction is more plausible, even when it says little about magnitude (Jones & Tukey, 2000). That defense collapses for publication bias, because the sign is known before we collect a single study. Selective reporting favors significant results; it does not run the other way. A test whose only defensible output is a direction we already know contributes little.

What it contributes instead is a verdict that is uninformative in both directions. A significant bias test conflates magnitude with detectability: in a large literature, a trivial and harmless amount of selection can reject the null. A nonsignificant bias test conflates small bias with low power: in a small literature, severe selection can easily fail to reach significance. Either way, the binary outcome tells us little about the quantity we care about. “There is publication bias, p < .05″ is, to borrow Cohen’s (1994) famous example, about as useful as “the earth is round, p < .05.” And “There was no evidence of publication bias, p > .05” is akin to “The earth is flat, p > .05.”

The deeper irony is that meta-psychologists have relocated the mistake they diagnosed. Original researchers treated significance as a discovery machine. Bias researchers sometimes treat significance as a bias-detection machine (Siegel et al., 2021). The error is identical: a difficult inferential problem is compressed into a binary decision.

Some have pushed the argument further, claiming that tests for publication bias are useless (e.g., Simonsohn, 2014). But the folly of nil-hypothesis testing, which incidentally undermines p-curve as much as many other significance-based methods, is not a reason to ignore publication bias. We do not abandon original research because the significance ritual is empty (Cohen, 1994). We replace the ritual with something more informative.

The reform for original research was to report effect-size estimates with confidence intervals that express uncertainty. The reform for bias detection should be the same. The goal is not to decide whether bias exists, but to estimate how much is present, how uncertain that estimate is, and whether the amount of bias consistent with the data changes the substantive conclusion.

Some publication-bias methods already estimate quantities of this kind, or carry the information needed to. Yet in practice that information is discarded, and the result is reduced to whether a test was significant or a method “detected bias” (Siegel et al., 2021). And no common metric for the amount of bias has been widely adopted.

The most natural metric is the excess of significant results. If a literature reports significant findings 80% of the time but the true probability of producing significant results is between 20% and 40%, we have clear evidence of substantial bias.

This is why the amount matters more than its presence. Bias can be easy to detect yet too small to change any conclusion, or large enough to overturn a conclusion yet impossible to detect in a small set of studies, where these tests have the least power (Renkewitz & Keiner, 2019). A binary test cannot tell these cases apart; an estimate with an interval can.

In short, meta-analysis needs the same methodological reform that original research needed. It is time to abandon the nil-hypothesis ritual and replace it with estimation: estimate the amount of publication bias, quantify the uncertainty with confidence intervals, and evaluate whether conclusions remain credible after adjusting for the plausible levels of selection.

Fortunately, unlike unpublished primary studies hidden in file drawers, the data behind published meta-analyses are often available or recoverable. That makes it possible to reexamine decades of meta-analytic conclusions and ask the question that matters: not whether publication bias can be detected, but whether the amount of bias compatible with the data changes what we should believe.

Simulations as Rhetorical Devices

The past decade has not been kind to experimental social psychology. Study after study failed to replicate and entire literatures have turned out to be built on nothing (a.k.a. statistical noise mining).

“Another day, another idol falls. This one has been teetering for years, so the collapse didn’t come as a shock. But that doesn’t make it any less painful.” (Michael Inzlicht).

It all started with a leading journal publishing an article with the crazy claim that people can foresee the future and practicing after a test can improve exam scores (Bem, 2011). This claim was quickly revealed to be false (and possibly a hoax, Gelman) after a big replication study failed to show the same results (Galak, J., LeBoeuf, R. A., Nelson, L. D., & Simmons, J. P., 2012).

In a media interview Bem explained that his experiments were never meant to be taken seriously. (Daniel Engber, 2017, Slate).

“If you looked at all my past experiments, they were always rhetorical devices. I gathered data to show how my point would be made. I used data as a point of persuasion, and I never really worried about, ‘Will this replicate or will this not?’

While the past decade has not been good for experimental social psychologists, it has produced a new group of psychologists to examine the causes of the replication crisis in experimental social psychology. As they look at the practices of research psychologists, they are meta-psychologists, psychologists who study other psychologists.

One of them is Blake McShane, who did his dissertation on statistical models to analyze time-series data (McShane, 2010). Given his background in statistics, managerial science, applied economics, and marketing, it is fair to say that he entered this field without first-hand experience of research practices that produced the replication crisis. He also does not cite seminal papers that foreshadowed the crisis by Cohen (1962, 1990, 1994). Instead, his main approach to examining meta-psychological questions appears to rely on his expertise in conducting simulation studies (McShane & Böckenholt, 2014, McShane, Böckenholt, & Hansen, 2016, 2020).

The problem with these simulation studies is that they repeat the same problems that plagued experimental social psychology at the meta-level. Just like Bem’s studies are not empirical tests, but rhetorical devices, McShane’s simulations are rhetorical devices to illustrate a point that does not require simulation evidence, namely.

[models] perform reasonably well in the setting for which they were designed, …[but] they are sensitive to deviations from their model assumptions.

In the 2016 article, the simulations violated assumptions of models that assume homogeneity and they failed. However, the simulations met the assumptions of another model and (no surprise) it worked well. However, McShane did not cite an earlier study that showed the model also has problems when its assumptions are not met (Hedges & Vevea, 1996).

Later simulation studies further confirmed that McShane’s preferred model does not work so well under realistic conditions (Carter et al., 2019), a finding not cited by McShane et al. in 2020. Pressed on this point that his simulations favored his preferred model, he might reply

“If you looked at all my past simulations, they were always rhetorical devices. I created conditions to show how things work when assumptions are met. I used simulations as a point of persuasion, and I never really worried about, ‘Does this apply to real data’ ”

In conclusion, a simulation that shows a model works when its assumptions are true and does not work when its assumptions are false is merely a demonstration, not an evaluation of a model under realistic conditions.

Who is Who in Social Psychology: Dolores Albarracin

“You can’t teach an old dog new tricks.” (Proverb)


Dolores Albarracín and the Defense of Old Social Psychology

Dolores Albarracín is a prominent social psychologist at the University of Pennsylvania whose work focuses on attitudes, persuasion, and behavior change. She has held major editorial positions in the field, including editor-in-chief of Psychological Bulletin from 2014 to 2020 and currently editor of the Attitudes and Social Cognition section of the Journal of Personality and Social Psychology (JPSP).

But which social psychology does she represent: the old social psychology of selectively publishing studies that confirmed researcher expectations, or the new open science that reports results independent of their desirability.

The answer can be found in two meta-analyses of the contested social (implicit) priming literature that has been the posterchild of the replication crisis. Albarracin published not one, but two meta-analyses in defense of social priming in Psychological Bulletin (Weingarten et al., 2016; Dai et al., 2023); the first one while she was editor of the journal.

The second one had to deal with the fact that many replication failures by new social psychologists willing to publish replication failures showed no evidence. Albarracin and her co-authors dismiss this evidence. They suggest that replication studies are themselves biased toward null results — a “reverse publication bias” — and therefore should be discounted or at least treated with the same suspicion as the old studies that used unscientific practices and selection of significant results to claim the effects are real and important.

The support for their claim is a blog-post about political bias in social psychology. In contrast, the publication bias in the older studies is not taken seriously leading to the dubious claim that implicit priming is a real phenomenon, even though Albarracin herself has not been able to demonstrate her own findings again in pre-registered new studies.

It is telling that somebody with this track-record and open hostility to the new and open social psychology is now editor of the very same journal that published Bem’s (2011) pseud-scientific evidence of extrasensory abilities. The irony is hard to miss. The journal that published false claims about extrasensory abilities is now controlled by somebody who makes false claims about open science practices and the credibility of implicit priming studies This is not a good look for social psychology in the 2020s.

Science is self-correcting, but nobody said that this process is fast and painless. It may require another decade for social psychology to fix all the problems that gave JPSP the name Journal of Pseudo-Scientific Psychology. Sadly, Albarracin is part of the problem, not of the solution. Fortunately, time is on the side of progress and the time for old social psychology is running out.

Open Science Requires Open Admission of Mistakes

Open Science in Psychology

What is open science? Isn’t open science a tautology like “new innovation.” If there is open science, what is closed science? The need for open science arises from the fact that many academic practices are unscientific. They benefit academics without advancing or even hurting science. For example, conducting experiments and not reporting the results when they do not show a favorable outcome is a common academic practice that many people would recognize as undermining science. In psychology, this academic practice is widespread and explains why psychology journals have success rates over 90% (Sterling et al., 1995). Aside from just not publishing unfavorable results, academics also use a number of questionable statistical practices to turn failures into successes (John et al., 2012). All of these practices are well known and accepted among academics who understand the pressure to publish, while the general public focuses on the outcome and not the personal consequences of individual researchers.

Open science is basically the idea of an utopia where academic work produces scientific progress and creates incentive structures that reward honest attempts to advance science rather than meeting invalid indicators like publication and citation counts that can be gamed and can waste millions of dollars without any real progress.

In psychology, Brian Nosek spearheaded the Open Science movement and founded the Open Science Foundation (OSF). He also wrote several influential articles to promote Open Science practices in psychology (e.g., Nosek & Bar-Anan, 2012; Nosek, Spies, & Motyl, 2012).

These articles laid out a comprehensive vision to reform unscientific and counterproductive practices and incentive structures in psychology. Key elements focussed on (a) aligning incentives so truth-seeking wins over career advancement, (b) restructuring the unit of research itself from small teams to distributed collaborations, and (III) promoting a culture of transparency, openness to criticism, and willingness to find out you were wrong.

The Open Science movement has changed psychology in ways that nobody in 2010 could have imagined. Helped by empirical evidence that many results in Brian Nosek’s field of social psychology could not be replicated (a replication rate of 25% in the Open Science Reproducibility Project, 2015), journals now often demand assurances that results are reported honestly and reward practices that limit researchers’ abilities to change hypotheses or results when the original results are disappointing.

However, in other ways, progress has been limited. The main problem is that open admission of mistakes is still rare and researchers fear that any admission of mistakes harms their reputation. Thus, the incentive structure continues to reward promoting false claims. This problem is exacerbated by psychological mechanisms that have been documented in psychological research for decades and are highly robust. Motivated biases make it easier for people to see mistakes in others’ work than in their own work. The Bible calls this “seeing a splinter in others’ eyes, but missing the beam in one’s own eye.” The Nobel Laureate Feynman warned fellow scientists, “The first principle is that you must not fool yourself — and you are the easiest person to fool.”

Motivated Blindness

Ironically Brian Nosek’s work on the IAT provides an example of motivated blindness. All his knowledge and intelligence that helped him to spot the problem in colleague’s work with small samples that does not replicate, does not help him to see the problems in his own work on implicit biases. Originally invented by Anthony Greenwald, Brian Nosek helped to promote the Implicit Association Test (IAT) as a measure of associations that are sometimes called implicit, automatic, or unconscious. The IAT is a reaction time task, but modern technology made it possible to administer it on a website, hosted by Project Implicit and backed by Harvard University.

The IAT was never validated to the psychometric standards required for individual assessment. In practice, it functions like a distorting mirror — reflecting back what people largely already know about their attitudes, buried under substantial measurement error. If it were presented that way, no one would object, and no one would need a warning. But Project Implicit does not present it that way. Instead, visitors are warned that the test may reveal something undesirable about themselves. That warning only makes sense if the results are trustworthy. A distorting mirror does not come with a warning — it comes with a laugh. By framing the IAT as capable of revealing uncomfortable truths, Project Implicit treats an unvalidated research tool as a diagnostic instrument.

The problem is that even in 2024, Brian Nosek is still unable to openly admit that the IAT does not measure implicit biases (reference) and that his own studies, which convinced him the IAT is valid, were flawed. For example, in one study he claimed that a weak correlation of r = .2 between racial bias on the IAT and self-reported racial attitudes demonstrated convergent validity (reference). This is false. A correlation of r = .5 between self-reported height and self-reported weight does not validate either measure — it simply shows that two different constructs share a common method. Convergent validity requires measuring the same construct with different methods, not different constructs with the same method. When the IAT is compared to other implicit measures, the correlations are equally weak and, more importantly, no higher than the correlations with self-report measures (Schimmack, 2021). The IAT therefore provides no evidence that it reveals something about individuals that they do not already know. If somebody is biased against a particular group, they know it. The IAT does not uncover hidden biases — it merely repackages what people can already report about themselves.

While Brian Nosek is no longer actively involved in IAT research, he is still associated with Project Implicit and has made no attempt to correct the misinformation about the IAT given to visitors of the website that even administers mental health IATs without proven validity. Moreover, his students continue to publish misleading articles that make false claims about the IAT. These articles are published in journals that claim to promote open science, but do not allow for open criticism of statistical errors in their publications.

The article “On the Relationship Between Indirect Measures of Black Versus White
Racial Attitudes and Discriminatory Outcomes: An Adversarial
Collaboration Using a Sample of White Americans” by Axt et al. (2026) seems to meet the latest standards of open science. The research team is diverse with different opinions about the IAT. Hypotheses are preregistered with a clear criterion to claim validity. Brain Nosek was not a collaborator, but strongly endorsed this article on social media as a posterchild of open science practices.

Yet, the paper had a major limitations. It totally ignored the criticism of earlier structural equation modeling studies that failed to take shared method variance into account (Schimmack, 2021) and it made the same mistake again. By including two IATs, the published model treated all shared variance between the two IATs as valid variance, ignoring the well known evidence that IAT scores are also influenced by factors unrelated to the associations being measured. The authors could have avoided this mistake because they inspected Modification Indeces that show problem with a theoretically specified model They used these modification indices to adjust the measurement model for self-ratings, but not for the two IATs.

This mistake itself is not the main problem. Even a large team of scientists can make mistakes, especially if they are not trained in psychometrics and are working with measurement models. The real problem is that the editor of the journal that published the article is unwilling to correct it (Schimmack, 2026). This decision does not meet Open Science standards of open admission of mistakes or even engagement with criticism. Open science requires open discussion and responding to scientific criticism. I emailed Dr. Axt on December 2nd about my concerns and reanalysis of his data, but did not receive a response. This reaction highlights how far we still have to come before we can reach Brian Nosek’s utopia of open criticism and open admission of mistakes. Marketing the IAT as a “window into the unconscious” (Banaji & Greenwald’s, 2013, words, not mine) was a mistake, but Greenwald, Banaji, and Nosek have yet to admit so openly. Instead, Project Implicit continues to give people invalid feedback and Harvard does not care. This is not Open Science. This is naked self-interest to preserve a reputation that was earned with the false promise of addressing racial bias in the United States of America.

Why Do I care?

After cognitive performance tests, the IAT is arguably the most influential psychological test. Implicit bias was a major topic during the 2016 presidential campaign. Hillary Clinton made implicit bias a campaign issue, claiming that many Americans still harbor implicit racial biases. Asked for comment, Greenwald relied on IAT results for the two candidates to “go out on a limb to predict that Clinton’s vote margin on November 8 will exceed the prediction of the final pre-election polls.” The opposite happened. Trump became president and created a new culture that made open expression of racial bias “great again.”

Greenwald’s trust in the IAT was not justified. The IAT had already failed to predict racial bias in the 2008 election that Barack Obama won despite widespread racial prejudice. The IAT did not predict this outcome, but self-reports showed that some people openly admitted to biases that predicted their voting intentions over and above party affiliation (Greenwald et al., 2009).

Hillary Clinton’s endorsement of implicit bias may have cost her votes. The notion of implicit bias is that white people no longer endorse racist ideology, are motivated to avoid racial biases, but are still unconsciously influenced by them. That narrative has not aged well. A decade later, a presidential candidate can stand on a debate stage and say “they’re eating the cats and dogs” to applause, and win. The problem America faces is not hidden bias operating below the threshold of awareness. It is open prejudice, stated plainly, rewarded electorally, and entirely accessible to self-report.

The implicit bias framework misjudged the landscape. It assumed that the social norm against racism was strong and stable, and that the remaining work was to address what operated beneath it. Instead, the norm itself collapsed. Many white Americans are fully aware of their racial biases, are not motivated to change them, and are willing to vote for a candidate who hesitated to distance himself from the KKK. These voters were probably more offended by the suggestion that they are motivated to be unbiased than by the accusation that they have racial biases. Implicit bias training — which cost organizations millions — failed to address the real problem because it was designed for a world in which people wanted to be fair but couldn’t help themselves. That is not the world we live in.

Conclusion

Open science promises to align academic structures, incentives, and practices with the scientific aim of discovering the truth. To do so, science needs to check itself, notice mistakes, and correct them. However, the incentive structure continues to work against this goal. It is telling that Brian Nosek, the most visible proponent of open science in psychology, is unable to follow his own open science principles and admit that his work on the IAT did not produce a valid measure of implicit biases.

One might think that Nosek is in an enviable position to admit past mistakes given his achievements in making psychology more open. He is the Executive Director of the Center for Open Science and has a legacy that does not depend on the IAT. Other psychologists, like John Bargh, built their careers on a single line of research. When social priming failed to replicate, there was little else to fall back on. Walking away from the IAT should be easier by comparison. The fact that Nosek is unable to acknowledge the problems of the IAT shows even more the power of motivated blindness. It also highlights the most important change that is needed to make psychology a science. We need to normalize failure and see it as the inevitable outcome of exploration. Every failure that is openly acknowledged is a learning opportunity that makes success more likely the next time. Daniel Kahneman is a rare example of a psychologist who admitted mistakes in public and gained in recognition as a result. Maybe we should give Brian Nosek a Nobel Prize for his open science work so that he can admit his mistakes about the IAT.

References

Axt, J. R., Connor, P., Hoogeveen, S., Clark, C. J., Vianello, M., Lahey, J. N., Hahn, A., To, J., Petty, R. E., Costello, T. H., Mitchell, G., Tetlock, P. E., & Uhlmann, E. L. (2026). On the relationship between indirect measures of Black versus White racial attitudes and discriminatory outcomes: An adversarial collaboration using a sample of White Americans. Journal of Personality and Social Psychology. Advance online publication. https://dx.doi.org/10.1037/pspa0000480

Greenwald, A. G., Smith, C. T., Sriram, N., Bar-Anan, Y., & Nosek, B. A. (2009). Implicit race attitudes predicted vote in the 2008 U.S. presidential election. Analyses of Social Issues and Public Policy, 9(1), 241–253.

Nosek, B. A., & Bar-Anan, Y. (2012). Scientific Utopia I: Opening Scientific Communication Psychological Inquiry, 23(3), 217–243. DOI: 10.1080/1047840X.2012.692215

Nosek, B. A., Spies, J. R., & Motyl, M. (2012). Scientific Utopia II: Restructuring Incentives and Practices to Promote Truth Over Publishability. Perspectives on Psychological Science, 7(6), 615–631. DOI: 10.1177/1745691612459058

Nosek, B. A. (2024, November 8). Highs and lows on the road out of the replication crisis [Interview]. Clearer Thinking with Spencer Greenberg, Episode 235.

Schimmack, U. (2021). The Implicit Association Test: A method in search of a construct. Perspectives on Psychological Science, 16(2), 396–414. https://doi.org/10.1177/1745691619863798

Schimmack, U. (2021). Invalid claims about the validity of Implicit Association Tests by prisoners of the implicit social-cognition paradigm. Perspectives on Psychological Science, 16(2), 435–442. DOI: 10.1177/1745691621991860


Closed Review == Censorship

Anonymous Closed Peer-Review is Censorship

Every self-interested entity in power wants to control public opinion. Billionaires buy newspapers, not to make more money, but to use their money to push their personal agenda. Totalitarian governments control access to free information to keep their citizens’ uninformed. The same human behavior is also visible in science, but it is often ignored.

British lords invented the “peer” (not you and me, but other lords) review system when they engaged in scientific debates as a hobby. Today, science is a billion dollar industry and scientists are self-interested actors in this system. Closed peer-review is still used to sell the public the impression that scientists control themselves to ensure that published articles meet the highest standards of scientific research. In reality, the closed peer-review system is used to control information and repress criticism.

The ability to influence the information that gets the stamp of peer-review approval is also the main motivation to take on the thankless job as an editor. The only reward is to decide which small number of submissions will get published or not. High rejection rates are used to claim rigorous quality control, but in reality, they give editors power to influence the narrative.

The problem is amplified at journals that focus on a specific narrow topic. These journals were often created by scientists who were not able to publish their work in other journals because their work was not considered important to the editors of those journals. For example, Cognition and Emotion was created in 1991 because psychology shunned research on emotions and even after the affective revolution in the 1980s, it was difficult to publish emoiton research in mainstream psychology journals.

Creating a journal to publish important work itself is a positive response to censorship. Rickard Carlson and I also used this approach to make it easier to publish research on meta-psychological topics that were difficult to publish elsewhere. However, the danger is that oppressed groups become oppressors, when they gain power. And closed peer-review gives editors at these new journals the power to control the narrative, just that it is now their narrative and their self-interests that decide what gets published. The only way to avoid this trap is to dismantle the power structure. That is what Rickard did with Meta-Psychology. First, articles are not rejected. They are improved until they meet basic scientific standards. Thus, there is no tool to suppress work because it is “not novel enough,” “only a small increment,” “outside of the scope of this journal,” or just a desk rejection with a note that the journal just cannot publish all of the important work that is done. The real reason is often that the editor did not like a paper.

In short, closed peer-review is not what the general public thinks it is. Rather than ensuring that research meets basic scientific standards, it is used to reward people to follow the party line and punish people who want to publish critical work.

Open Science Reforms

In psychology, the academic discipline I know because I worked in it for over 30 years now, the problem of censorship became apparent during the replication crisis in the 2010s. Peer-review had failed to ensure that published results are scientifically valid. Lack of training and understanding of science itself was partly to blame, but the bigger reason was that peer-reviewers were happy to publish bad research because they were doing the same bad research and were interested in publishing these results that benefited their own work. Yes, I am talking about the implicit revolution (Greenwald’s words, not mine) that seemed to show that much of human behavior was caused by mindless responses to situational cues without even noticing it. Call it implicit, automatic, or unconscious, experiment after experiment seemed to support these claims. In reality, research on the unconscious worked very much like Freud’s model of unconscious process. Undesirable results were repressed and only results that showed support for researcher’s claims were published. This became apparent after Bem even showed time reversed unconscious processes, which nobody was willing to believe. When other studies were replicated, they also failed to provide support for other claims and the implicit revolution imploded. Peer-review had failed as a quality control mechanism. Rather censorship had created a bubble of false findings. It doesn’t take a psychoanalyst to realize that the realization was painful and that many old researches resorted to defense mechanisms to avoid the emotional consequences of realizing that their achievements were illusory.

Open science requires open sharing of all findings and arguments. It also requires that conclusions are consistent with the evidence and logically consistent. This open exchange cannot happen in a closed peer-review system where editors control the narrative. The new quality assurance is not “peer-reviewed,” but “open peer reviews,” and publication of all arguments on both sides. It is also important to get rid of journal rankings to evaluate the quality of research. Journal rankings only ensure that editors of prestigious journals have even more power to control the narrative. I experienced this first hand. When I submitted my first critique of the Implicit Association Test to the prestigious journal “Perspectives on Psychological Science,” the editor rejected it. When I tried again several years later, a new editor accepted it. Neither decision was based on the quality of the work or the argument, it was just a personal preference.

A Scientific Utopia

Most editors also do not read articles they handle or provide their own comments. The bias is often introduced by picking reviewers that will like or dislike a paper (I know, I was Ed Diener’s henchmen, his words, not mine). So, they really do not add anything of value. Even current AI (large language models) are better able to evaluate the scientific merits of a paper and we can replace human editors with AI, a faster, more cost effective, and less biased way to make decisions about publications that are essential for young scientists’ careers.

Erratum: More concerns about the z-curve method

Scientific progress has been slow because humans are not disinterested processors of information. Once they have concluded that some belief is true, their information processing is biased towards verifying that truth rather than looking for disconfirming evidence.

Willful ignorance is the selective processing of confirmatory information and the avoidance of sources that may expose the believer to contradictory information. However, sometimes challenging information is unavoidable. Scientists who want to publish their work are constantly exposed to negative comments. When confronted with criticism, there are a number of strategies that serve different purposes. A constructive response examines the validity of the criticism, responds to valid concerns, adjusts claims accordingly, and may still make a useful contribution. A defensive response to valid criticism engages in pseudo-scientific arguments that avoid the key concern and leads to an unproductive exchange that cannot have a resolution because the goal is to maintain a false belief.

While critics initiate a discussion about potential errors, the roles are not fixed. Once the criticism is made, the person criticized responds to it and may find errors in the critic’s arguments. Now the roles are reversed and the critic may respond to this criticism in defensive ways, accusing the person being criticized of being defensive. This exchange quickly deteriorates into a childish exchange of shouting “I am right. You are wrong” at each other. A more mature response is to allow for errors being made on both sides and carefully examine the arguments. This is the aim of my response to Erik van Zwet’s second blog post about z-curve, “More concerns about z-curve.

The Substance

In this second post, Erik reports one new simulation scenario. In that scenario, he points to two problems. The main criticism is that the confidence interval for the Expected Discovery Rate (EDR) does not achieve its nominal 95% coverage. The second concern is that the confidence interval for the null-component weight can collapse to zero width, which he interprets as a sign of instability or misspecification in the internal mixture fit.

The second point is the less important one. Z-curve is a finite-mixture model that approximates the distribution of test statistics using weights on several discrete components. It is well understood that these component weights are not themselves substantively meaningful parameters when the true data-generating process is continuous. Different mixtures can yield nearly identical estimates of the quantities z-curve is designed to recover. For that reason, poor coverage of confidence intervals for individual component weights is not, by itself, a serious problem. In particular, the weight of the zero component is not used in z-curve the way a null-component weight is used in models that directly estimate false positive rates. These intervals appear in the output, but they are not the primary inferential target.

What matters is coverage for the main estimands: the Expected Replication Rate (ERR) and the Expected Discovery Rate (EDR). Erik does not mention that the ERR interval appears to perform adequately in this scenario. Thus, the central substantive criticism is narrower: in this particular simulation setting, the EDR confidence interval appears to undercover.

The Response

The specific scenario assumed that all studies had the same power, which implies not only the same sample size, but also the same population effect size. Brunner and Schimmack (2020) already noted that z-curve can have problems in this situation when the true noncentrality parameter falls between two default components. That is exactly Erik’s scenario: mean power is 32%, corresponding to z = 1.5, midway between the default components at z = 1 and z = 2.

Brunner and Schimmack (2020) did not emphasize this problem because most real datasets show substantial heterogeneity in sample sizes and effect sizes (van Erp et al., 2017). Even direct replications of the same paradigm across labs vary in effect size (Klein et al., 2017). Thus, Erik’s critique is based on a known difficult case for z-curve.2.0, but not one that resembles most real applications.Use this instead:

To address this valid concern, z-curve 3.0 was revised to first test for very low heterogeneity. When the data appear unusually homogeneous, the model estimates where a single component would best fit the distribution and then shifts the default grid so that one component is centered near that value. In Erik’s scenario, this places a component near z = 1.5 instead of forcing the fit to choose between z = 1 and z = 2.

The new results are therefore limited to Erik’s specific concern: whether z-curve.2.0 provides adequate coverage for homogeneous data when the true noncentrality parameter falls between two default components.

I validated z-curve 3.0 with the standard simulation code that was used to validate z-curve.2.0 in the Uli simulation design. These simulations across 192 scenarios were validated with just 50 significant results to produce coverage over 95% in most scenarios. To simulate a non-centrality parameter of z = 1.5, I used a standardized mean difference of d = .30 and a sample size of N = 100 (.3 / (2/sqrt(100) = 1.5) . Figure 1 shows the results for 50,000 significant results. Z-curve is able to predict the distribution of the non-significant results based on the model fitted to the significant results well and the estimates of EDR and ERR are accurate and the confidence intervals are tight.

Coverage for the ERR and EDR estimates was tested with k = 50, 500, 5,000, and 50,000. All simulations showed coverage over 95% (Results). In short, z-curve.3.0 now also performs well with homogenous data and can do so quickly with the density method.

In sum, Erik noted that the default method of z-curve.2.0 fails to produce adequate confidence intervals for the EDR estimate in one simulation with homogenous data and a non-centrality parameter between two default components. I responded to this valid criticism by improving z-curve. Z-curve.3.0 now handles homogeneity and heterogeneity in power well and provides credible confidence intervals.

In the comment section Erik writes. “Indeed, as I wrote: “Note that I’m violating the assumption of the z-curve method, but in a way that would be difficult to detect from limited data. That’s the point: You can fix this by changing the default “mu grid”, but you wouldn’t know that.”

As I showed here, this statement is an error. It is very easy to diagnose the problem by estimating the heterogeneity of the data and then adjust the grid according to a preliminary model that is more consistent with the data. The ability of z-curve.3.0 to work in this scenario shows that the problem is fixable. Thus, Erik’s criticism is invalidated by the evidence. Any new evaluations of the z-curve method need to examine the performance of z-curve.3.0.

Learning from van Zwet’s Concerns: Z-curve.3.0

Unwillingness to Admit Mistakes

It’s sooooo frustrating when people get things wrong, the mistake is explained to them, and they still don’t make the correction or take the opportunity to learn from their mistakes.

This could have been written by me or many other people who are in the business of calling out other people’s mistakes. In theory, that would be all scientists because science is supposed to progress by correcting mistakes. However, academia is not science and many academics don’t like to face their own mistakes. The more their status and reputation depends on some claim they made in the past, the more reluctant people are to admit that they were wrong. Max Plank famously declared that science only progresses when pig-headed prominent scientists die and the field can move on. But humans are human and public admission of mistakes is not a virtue in modern capitalist science that reward self-promotion and sexed-up research findings.

Anyhow, I digress. The quote is from Gelman’s blog post about “Learning from mistakes (my online talk for the American Statistical Association, 2:30pm Tues 30 Jan 2024)

While it is true that the incentives are against public admission of mistakes, there are notable exceptions. Daniel Kahneman, after he won a Nobel Prize, was able to admit some mistakes. Maybe it takes a Nobel to overcome nagging feelings of self-doubt and defensiveness. I hope not. I have corrected some of my mistakes, but I have to admit, that it sometimes took a long time to admit them. At the same time, I have also pushed back against critics who were wrong. The real problem is of course to know the difference. Accept valid criticism, reject invalid criticism, requires knowing what is valid and what is invalid. Thus, the requestion for all actors, critic, person being criticized, and observers is “Who is right?”

How To Respond to Valid and Invalid Criticism

In another blog post, Gelman gives advice to people who have been criticized about better or worse ways to respond to criticism (A ladder of responses to criticism, from the most responsible to the most destructive | Statistical Modeling, Causal Inference, and Social Science)

The content of the blog post, however, conflates responding to criticism with responding to an error in one’s work.

Consider the following range of responses to an outsider pointing out an error in your published work:

  1. Look into the issue and, if you find there really was an error, fix it publicly and thank the person who told you about it.
  2. Look into the issue and, if you find there really was an error, quietly fix it without acknowledging you’ve ever made a mistake.
  3. Look into the issue and, if you find there really was an error, don’t ever acknowledge or fix it, but be careful to avoid this error in your future work.
  4. Avoid looking into the question, ignore the possible error, act as if it had never happened, and keep making the same mistake over and over.
  5. If forced to acknowledge the potential error, actively minimize its importance, perhaps throwing in an “everybody does it” defense.
  6. Attempt to patch the error by misrepresenting what you’ve written, introducing additional errors in an attempt to protect your original claim.
  7. Attack the messenger: attempt to smear the people who pointed out the error in your work, lie about them, and enlist your friends in the attack.

As you can see, there is no option to look at the issue, find a mistake in the criticism, point out the mistake, and the critic apologizes and thanks the person being criticized for engaging constructively and taking time to address their concern.

A Case Study

Taken, Erik van Zwet’s post “Concerns about z-curve “as an example. The post contains several mistakes about z-curve. Some mistakes are glaring, like being a reviewer of z-curve and then claiming it was not vetted by experts.

1. Gelman had made the sweeping claim that many statistical tools are “never vetted by experts, and often are just “verified” by a few simulations.” van Zwet then writes “I believe that another meta-analytic method called z-curve (Brunner and Schimmack (2020)Bartos and Schimmack (2022)Schimmack and Bartos (2023)) has similar problems ”

The strange fact, not mentioned by van Zwet on his blog post, is that he wrote a favorable review of z-curve when he was a reviewer of z-curve.2.0. Claiming that z-curve was not reviewed by experts implies that he is not an expert, but if he is not an expert, it undermines his critique of z-curve.

2. van Zwet then claims that the z-curve method is based on the assumption that the absolute values of the SNRs have a discrete distribution supported on 0,1,2,…, 6.  That statement confuses the default settings of the z-curve package with the z-curve method. Criticizing these defaults is fine, but confusing default settings and a method is not. Especially Bayesian statisticians like Gelman and van Zwet should know the difference.

If somebody uses Gelman’s statistical tool, stan, with bad priors, it leads to bad results. The problem is not the tool, but the prior. I have made this point clear in the comment section and pointed out that z-curve handles some specific edge-cases where the defaults fail well by changing the defaults.

3. In the conclusion, van Zwet makes generalizes from a single scenario that shows z-curve underestimates uncertainty to imply that z-curve is always unreliable. “In my opinion, statistical methods should be reliable when their assumptions are met. I don’t think unreliable methods should be used because no better methods are available.”

Once again, this is like saying nobody should use Gelman’s stan program to analyze data because one application resulted in a false conclusion. Non-sensical, unscientific, and clearly a mistake that only Reviewer B would make because the goal is not to advance science, but to be a nasty reviewer for reasons that remain unknown (e.g., sexual frustration, grant application failed, realizing that academia is a waste of time, no hobby, etc.).

How I respond to valid criticism

Let me show how I respond to valid concerns. Yes, in the specific scenario picked by van Zwet, z-curve.2.0 was overconfident and produced confidence intervals that were too narrow and missed the true value more often than a 95% confidence interval should, namely more than 5 out of 100 times. That is a valid criticism of z-curve.2.0.

I was already working on improving z-curve. Using van Zwet’s scenario, I was able to use information in the data to alert z-curve to scenarios that provide little information about the expected discovery rate (van Zwet’s own simulation had 40% data that contained absolutely no information). I tested z-curve.3.0 with van Zwet’s scenario and 99 out of 100 simulations contained the true value. Thus, the new confidence intervals provide accurate information about lack of information about the EDR in the data.

Of course, z-curve is not magic. As the plot shows, the EDR is an estimate of the distribution of non-significant results based on only the significant results. When there are few informative z-values just below significance (z = 1.96 to 2.96), the EDR cannot be estimated. Z-curve.3.0 realizes this and gives a wide confidence interval that ranges from 15% to 98%. This is informative because it tells users that the EDR cannot be estimated and the point estimate cannot be trusted. However, the confidence interval will be smaller and more informative in other situations and with larger sets of studies.

In short: z-curve.2.0 is dead. Long live z-curve.3.0

Now, this is how you respond to valid concerns and demonstration of errors. You learn from them and fix them. That is how real science advances and z-curve has been developed, evaluated, and improved for over 10 years now.

Waiting for Gelman and van Zwet’s Response to this Criticism

It will be interesting to see how van Zwet and Gelman respond to this criticism of their criticism. The ladder of responses is clear and now also includes pointing out errors in my response or in z-curve.3.0 In the age of preregistration, let me preregister my prediction.

4. Avoid looking into the question, ignore the possible error, act as if it had never happened, and keep making the same mistake over and over.

I hope this is a mistake that I am happy to correct when proven wrong.

The Gelman Prior: Don’t Trust Anything

Andrew Gelman is a statistician who is working for Columbia University. He also maintains a blog post where he shares his opinions about many topics, including the replication crisis in psychology and related fields like behavioral economics. He is not an expert in either field, but that does not prevent him from evaluating the research in these areas. But you do not have to read a specific blog post by him because the result is often the same. The research is not credible, sample sizes are too small, studies are selected for significance, and meta-analyses are not trustworthy. In his favorite area of statistics that uses prior assumptions to make sense of actual data, this is known as a dogmatic prior. No amount of data will reverse the conclusion that is already implied by a dogmatic prior. So, you really do not need data.

As you may have guessed, I don’t like the guy. I think he is a jerk, and that may cloud my evaluation of him. However, I do have data to support my claim that the Gelman’s statements often reflect his prior assumptions and are immune to data. He says so himself on his blog post.

After discussing some problems with a meta-analysis of nudging studies (a Nobel prize winning idea in behavioral economics), Gelman writes:

Just to be clear: I would not believe the results of this meta-analysis even if it did not include any of the above 12 papers, as I don’t see any good reason to trust the individual studies that went into the meta-analysis. It’s a whole literature of noisy data, small sample sizes, and selection on statistical significance, hence massive overestimates of effect sizes.

What are small sample sizes (some of these studies have hundreds of participants)? Where is the evidence that selection leads to MASSIVE overestimation. Gelman has no answers to such scientific questions about the evidence because he does not care about the data. His prior is sufficient to dismiss an entire literature, not just a few bad studies.

Did I cheery-pick this example? Should you trust me? To find an answer to these questions you can use AI that can read Gelman’s blog within seconds. Share one of his blog posts where he reversed a prior belief in response to empirical data. I am waiting.

The problem is not that Gelman is opinionated and shares his opinions on a blog (some people may say that is also true of myself). The problem is that he has blind followers that seem to confuse believing Gelman’s opinions with meta-science. Actual understanding of problems in science requires investigating these problems with empirical methods and draw conclusions from data; not believing in conclusions that rest on unproven assumptions.

Countering Misinformation about Z-Curve on Gelman’s Blog

I am all in favor of open science and a critic of closed pre-publication peer-review. The downside of open communication is that there is no quality control and internet searches will amplify misinformation. This is the case with Erik van Zwet’s critique of z-curve. Even though I addressed his criticisms in the comment section, search engines – like humans – do not scroll to the end and process all information. I have even addressed concerns about z-curve.2.0 by improving z-curve 3.0 to handle edge cases like the one used by van Zwet to cast doubt about z-curves performance in general. In science, facts trump visibility Z-curve.has been validated with many simulations across a wide range of scenarios and works well even with just 50 significant z-values. For more information, check out the Replication Index blog or the FAQ about z-curve page.

The bias in the Bing (AI) summary is evident when we compare it to Google search summary. Still makes a false claim about assumptions based on Erik van Zwet’s blog bost, but also avoids the dismissal of a method based on a single edge case that was easy to address and is no longer of concern in the new z-curve.3.0. In short, don’t trust the first generic response of AI. Use AI to probe arguments.

Google