RETRACTED: Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. 


In 2012, I published an article in Psychological Methods titled “The Ironic Effect of Significant Results on the Credibility of Multiple-Study Articles” (Schimmack, 2012). The paper introduced a simple but powerful idea: if a set of published studies reports a higher proportion of significant results than would be expected based on the estimated average post-hoc power, this suggests the results may be “too good to be true.” That is, the success rate of these studies may reflect publication bias, p-hacking, or other questionable research practices (QRPs) rather than honest science.

This logic, based on Sterling et al., 1995, has been foundational to several developments in meta-science and replicability research over the past decade. It also aligns with statistical common sense: if the probability of success is low, and we still see nearly universal success, something about the process likely isn’t transparent or unbiased.

But in 2024, Psychological Methods published an article by Pek et al. that challenges this logic at its core. They argue that using observed outcomes to evaluate the expected success rate — and then inferring bias when the observed rate is too high — is not just statistically questionable. They call it an ontological error: a fundamental mistake in the nature of inference itself, because (they claim) we cannot assign probabilities to events that have already occurred.

I disagree with this argument. Like most statisticians and meta-scientists, I believe that statistical inference is inherently about comparing observed outcomes to expectations under a model. That’s what a p-value does. That’s what every goodness-of-fit test does. That’s what replication rates and power estimates are meant to assess. The logic of comparing the observed to the expected is the backbone of empirical science — not a metaphysical error.

However, in the hypothetical world in which Pek et al. are correct, and more importantly, in a world where Psychological Methods treats their position as settled and unchallenged, my 2012 article becomes indefensible on its own terms. If it is indeed an ontological error to compare observed success rates to expected ones, then my article’s entire logic — and its main conclusion — are invalid.

We submitted a commentary in 2025 defending the logic of my 2012 paper and challenging the categorical claim made by Pek et al. That commentary was rejected without invitation to revise, indicating that the editor — and thus the journal — considers the matter settled. No debate, no dialogue. Just closure.

So, I’ve taken the logical next step. If the journal has adopted the position that the logic of my article is invalid at the most fundamental level, then the appropriate action is retraction. Not correction. Not a commentary. Retraction.

Below is the letter I sent to Fred Oswald, Editor of Psychological Methods, on July 23, 2025:


Subject: Formal Request to Retract Published Article
To: Fred Oswald, Editor, Psychological Methods
From: Dr. Ulrich Schimmack
Date: July 23, 2025

Dear Fred Oswald,

I am writing to formally request the retraction of the following article:

Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17(4), 551–566. https://doi.org/10.1037/a0029487

This article argued that the credibility of multiple-study papers can be evaluated by comparing the observed rate of significant results to the rate expected based on average post-hoc power. It concluded that when the observed success rate substantially exceeds the expected rate, the results are “too good to be true,” suggesting publication bias or questionable research practices.

In 2024, Psychological Methods published an article by Pek et al. that characterized this type of inferential reasoning as a fundamental ontological error. Specifically, they argue that statistical inferences based on observed outcomes are invalid because probabilities cannot be assigned to events that have already occurred. According to their reasoning, the method used in the 2012 article is not just flawed but conceptually incoherent.

A commentary was submitted to the journal defending the validity of this approach and challenging the ontological framing advanced by Pek et al. That commentary was rejected without invitation for revision, signaling that the journal considers the matter settled and sides with Pek et al.’s position.

If Psychological Methods now endorses the view that the central logic of the 2012 article is invalid, then the article no longer meets the journal’s standards for methodological soundness. Its continued presence in the scholarly record, without correction or rebuttal, misleads readers into believing that its conclusions are supported by valid statistical reasoning.

Therefore, I request the retraction of the article. This is not a concession of error by the author, but the necessary course of action given the journal’s editorial stance. To hold that the logic is invalid yet leave the article uncorrected would be editorially inconsistent and undermine the integrity of the journal’s standards.

Sincerely,
Ulrich Schimmack
Department of Psychology
University of Toronto Mississauga


If the journal wants to stand by Pek et al.’s claim that comparing observed outcomes to expected frequencies is a category error, then it must have the editorial courage to follow through: retract past articles that used that logic.

If, instead, the journal is unwilling to retract, then it implicitly concedes that the issue is not settled and deserves ongoing debate. In that case, our rejected commentary should be reconsidered for publication.

Leave a Reply