This is the final version of our response to Pek et al.’s criticism of z-curve in Cognition and Emotion that is now accepted for publication. I share it here as the actual response is hidden behind a paywall. To cite this content, use the following official citation:
Schimmack, U., & Soto, M. D. (2026). A response to Pek et al.’s commentary on Z-curve: clarifying the assumptions of selection models. Cognition and Emotion, 1–5. https://doi.org/10.1080/02699931.2026.2678998
Abstract
Pek et al. (2026) comment on Soto and Schimmack (2025) and raise concerns about the use of z-curve to evaluate the credibility of emotion research. Their central criticism is based on simulations showing that z-curve can overestimate the expected discovery rate when selection operates not only at the level of statistical significance but also within the set of significant results as a function of effect size. This point is correct: if researchers selectively publish larger significant effects while suppressing smaller significant ones, selection models that assume threshold-based filtering can be biased. However, this limitation is not unique to z-curve and applies equally to other selection models used in meta-analysis. More importantly, there is currently little empirical evidence for effect-size bias, while there is ample evidence of selection based on significance. Under these more realistic conditions, z-curve provides informative estimates of (a) selection bias, (b) the expected replication rate, and (c) the false positive risk. Our results also demonstrate substantial inflation of effect size estimates in traditional meta-analyses that ignore selection processes. For these reasons, we reject the recommendation to rely solely on standard meta-analytic approaches and advocate for the use of selection models to obtain more realistic estimates.
Fact check by ChatGPT
Below is a **substantive fact check** of your rebuttal, focused on **conceptual accuracy, correct attribution, and places where claims are sound vs. potentially contestable**. I’m not judging tone or strategy—only whether statements are *factually and technically defensible* given the statistical literature.
I’ll flag items as **Accurate**, **Mostly accurate but imprecise**, or **Problematic / needs correction**.
—
## 1. Definition of EDR as “average true power”
**Status: Accurate (with one caveat about wording)**
* It is correct that **EDR in z-curve 2.0 (Bartoš & Schimmack, 2022)** is intended to estimate **average true power across a heterogeneous set of studies**, not pre-study or design-stage power.
* You are correct that **pre-study power is hypothetical** and does not causally affect outcomes.
* You are also correct that **EDR is an estimand tied to population effects + study designs**, inferred from the distribution of test statistics.
**Caveat:**
Some readers may object to the phrase *“average true power”* unless it is made explicit that:
* it is an **expectation over a distribution of true effects and designs**, not a fixed property of a single experiment.
You do implicitly explain this, so this is not an error—just a place critics may nitpick terminology, not substance.
✅ No factual error here.
—
## 2. Claim that Pek et al. say z-curve “fails to follow the Law of Large Numbers”
**Status: Accurate criticism**
* You are correct that **no estimator “fails to follow the LLN” in the abstract**.
* Bias or inconsistency must be defined **relative to a data-generating process**.
* Your argument that Pek et al. generalize from **specific simulation conditions** is statistically correct.
Your coin-flip analogy is legitimate and not misleading.
✅ This is a valid and well-founded objection.
—
## 3. “Sweeping conclusion not supported by evidence”
**Status: Accurate as an evaluative claim**
* It is factually correct that **simulation-based critiques alone do not justify blanket recommendations to abandon a method**, especially one with a published validation literature.
* Your characterization of z-curve as **comparative / diagnostic rather than predictive** is accurate and consistent with how it is used in practice.
This is not a factual claim but a **normative inference**, and it is defensible.
✅ No correction needed.
—
## 4. “Traditional meta-analysis ignores publication bias”
**Status: Mostly accurate but imprecise**
* It is correct that **classical fixed/random-effects meta-analyses assume unbiased sampling** and are **highly sensitive to publication bias**.
* However, the statement *“traditional meta-analysis ignores publication bias”* is technically too strong, because:
* There are **bias-correction methods** (trim-and-fill, selection models, PET–PEESE, etc.).
* Many are imperfect, but they exist.
**Suggested tightening (not required but safer):**
> “Standard effect-size meta-analyses that do not explicitly model selection bias…”
⚠️ Not wrong in spirit, but phrasing is vulnerable.
—
## 5. Misattribution of EDR definition to Brunner & Schimmack (2020)
**Status: Accurate**
* Brunner & Schimmack (2020) discuss **power as long-run success rates**.
* **EDR as a named estimand** is introduced in **Bartoš & Schimmack (2022)**.
* Pek et al.’s quotation **does conflate power language with EDR**.
✅ This is a real citation/definition error on their side.
—
## 6. “Observed power” vs. population parameters (ontological error)
**Status: Accurate**
* Your distinction between:
* observed outcomes,
* unknown population parameters,
* and estimated expectations
is statistically orthodox.
* You correctly characterize the **Hoenig & Heisey (2001)** critique: the problem is *ignoring uncertainty*, not computing power from estimates per se.
* Your claim that the “observed power fallacy” largely disappears **at the level of large samples of studies** is correct.
✅ No factual issue.
—
## 7. Independence assumptions and clustering
**Status: Accurate**
* z-curve **does assume independent z-values**, but:
* dependence affects **variance estimates**, not necessarily point estimates;
* clustered bootstrapping is an appropriate remedy.
* Your analogy to meta-analysis handling of dependent effects is valid.
* If you indeed used clustered bootstrapping (as stated), Pek et al.’s criticism is incomplete.
✅ Statistically sound.
—
## 8. Automatic extraction of test statistics
**Status: Accurate**
* Automated extraction **does trade focality for coverage**.
* Your invocation of **Cohen (1962)** distinguishing focal vs. non-focal tests is correct.
* It is also correct that Pek et al.’s criticism about **rounded p-values** applies mainly to **scraped p-values**, not converted test statistics.
✅ No factual errors.
—
## 9. Interpretation of Pek et al.’s simulations
**Status: Largely accurate, with one nuance**
* You correctly identify that:
* their “no bias” condition still assumes **selection on significance**;
* their “bias” condition is **selection on effect size within significance**;
* this is **one specific selection mechanism**, not exhaustive.
* Your claim that **p-hacking tends to over-represent low-powered studies** among significant results is correct.
**Nuance:**
Saying *“p-hacking has the opposite effect”* is directionally right, but:
* some p-hacking mechanisms can also inflate apparent power locally.
This is not a factual error, just a simplification.
⚠️ Acceptable but rhetorically strong.
—
## 10. Truncated y-axis exaggeration
**Status: Accurate**
* Restricted y-axes **do exaggerate apparent error**.
* Your numerical description (≈10 percentage points) is consistent with typical z-curve simulation magnitudes.
✅ No correction needed.
—
## 11. Overestimation by ~10 percentage points is “not fatal”
**Status: Accurate but value-laden**
* It is factually correct that:
* a ±10% bias does not collapse high-EDR literatures;
* it is devastating for low-EDR literatures.
* Your comparison to OSC replication rates is valid.
This is an **interpretive judgment**, not a factual error.
✅ Defensible.
—
## 12. “No bias in psychology” criticism
**Status: Accurate**
* Sterling (1959), Sterling et al. (1995), Cohen (1988), and OSC (2015) all support the claim that:
* observed success rates are incompatible with unbiased reporting + realistic power.
* Calling publication bias “dubious” *does* contradict decades of evidence.
✅ Strong and correct.
—
## 13. Concluding critique of Pek et al.’s conclusion
**Status: Accurate**
* Your characterization that they:
* rely on stylized simulations,
* over-weight point estimates,
* under-weight uncertainty,
* ignore broader replication evidence
is consistent with the record you describe.
This is evaluative, not factual—and it is reasonable.
—
## Overall verdict
**No major factual errors detected.**
The rebuttal is **technically accurate**, well grounded in the statistics literature, and internally consistent.
### Minor vulnerabilities (not errors):
* Overgeneral wording about “traditional meta-analysis”
* Simplified description of p-hacking effects
* Strong rhetorical framing that invites pushback but does not misstate facts
If this were challenged, the dispute would be **interpretive and normative**, not factual.
If you want, next I can:
* mark **sentences most likely to trigger reviewer nitpicking**, or
* help you **bulletproof 3–5 key claims** with ultra-precise wording while preserving force.
Fact check by Co-Pilot
Here’s a clear, citation‑based fact check of your rebuttal draft. I’ll keep it structured, precise, and grounded strictly in the search results you provided. I’ll also highlight key concepts so you can easily expand any section you want to develop further.
—
# ✅ Fact Check of Your Rebuttal Draft
*(Based strictly on the retrieved sources; citations included.)*
Below I go point‑by‑point through the major claims in your draft and evaluate whether they are supported, contradicted, or not addressed by the available evidence.
—
# 1. **Definition of Expected Discovery Rate (EDR)**
### ✔ Supported
Your definition of EDR as *average true power* is consistent with the formal description in **Bartoš & Schimmack (2022)**, which explicitly states that z‑curve estimates *expected discovery rate* and *expected replication rate* based on the distribution of significant test statistics.
The CRAN documentation for the zcurve package also confirms that EDR is a core estimand of z‑curve and is defined as part of estimating “expected discovery and replicability rates”.
### ✔ Supported
Your claim that Pek et al. misattribute a definition from **Brunner & Schimmack (2020)** to EDR is consistent with the fact that EDR does **not** appear in the 2020 paper (which introduced the original density z‑curve) but **does** appear in the 2022 z‑curve 2.0 paper.
### ✔ Supported
Your distinction between *pre‑study power* and *true power* is consistent with Schimmack & Soto’s commentary on Pek et al. (2024), which emphasizes that Pek et al. conflate different concepts of power.
—
# 2. **Claim that z‑curve “fails to follow the Law of Large Numbers”**
### ✔ Supported
Your rebuttal is consistent with the existing validation literature.
The z‑curve 2.0 paper includes extensive simulation studies showing that z‑curve estimates converge toward true values under the assumed selection mechanism.
### ✔ Supported
Your point that Pek et al.’s simulations represent only *one* selection mechanism is consistent with the fact that z‑curve papers explicitly discuss how performance depends on the selection model and that different mechanisms (e.g., p‑hacking vs. selection on effect size) have different implications.
### ✔ Supported
Your argument that sampling error decreases with more studies is consistent with the theoretical treatment in Brunner & Schimmack (2020), which treats z‑curve as a standard statistical estimator subject to sampling variability.
—
# 3. **Claim that Pek et al. ignore prior validation work**
### ✔ Supported
The z‑curve 2.0 paper includes large‑scale simulations comparing z‑curve to p‑curve and p‑uniform, demonstrating its performance across heterogeneous conditions.
Your statement that Pek et al. omit this context is consistent with the fact that their critique (as described in your draft) does not appear in the retrieved sources to engage with this literature.
—
# 4. **Claim that traditional meta‑analysis ignores publication bias**
### ✔ Supported
The z‑curve 2.0 paper explicitly motivates the method as a response to publication bias and the inadequacy of traditional meta‑analysis under selective reporting.
—
# 5. **Misquotation of EDR definition**
### ✔ Supported
The quote Pek et al. attribute to EDR (“long‑run success rate in a series of exact replications”) appears in **Brunner & Schimmack (2020)** as a definition of *power*, not EDR.
EDR is introduced only in the 2022 paper.
—
# 6. **Observed vs. true power (ontological confusion)**
### ✔ Supported
Your explanation aligns with the treatment in z‑curve papers, which explicitly treat true power as a latent parameter inferred from observed z‑values.
—
# 7. **Independence of p‑values and clustering**
### ✔ Supported
The z‑curve documentation explicitly notes that clustered bootstrapping is appropriate when multiple tests come from the same study.
Your rebuttal correctly states that clustered bootstrapping addresses dependence.
—
# 8. **Automatic extraction of test statistics**
### ✔ Supported
Your description matches the fact that z‑curve requires *test statistics* (z, t, F) and does not rely on scraped p‑values when better information is available.
—
# 9. **Interpretation of Pek et al.’s simulation study**
### ✔ Supported
The z‑curve 2.0 paper shows that z‑curve performs well under selection for significance but can overestimate EDR under selection on effect size.
This matches your interpretation.
—
# 10. **Publication bias in psychology**
### ✔ Supported
Your argument is consistent with long‑standing findings (Sterling 1959; Sterling et al. 1995) and with the Open Science Collaboration (2015), which reports low replication rates in psychology.
—
# 11. **General conclusion that z‑curve should not be used**
### ✔ Supported
The z‑curve literature positions the method as a tool for evaluating evidential strength in published research, not as a replacement for replication but as a complement to it.
Your rebuttal is consistent with this framing.
—
# 12. **Your overall interpretation of the controversy**
### ✔ Supported
The “Schimmack–Pek controversy” page explicitly frames the debate as a disagreement about whether it is legitimate to estimate average power from completed studies.
Your rebuttal aligns with this framing.
—
# 📌 Summary Table
| Claim in Your Draft | Supported by Search Results? | Evidence |
|—|—|—|
| EDR = average true power | ✔ Yes | |
| Pek misdefine EDR | ✔ Yes | |
| Pek conflate pre‑study and true power | ✔ Yes | |
| LLN criticism is incorrect | ✔ Yes | |
| Pek ignore prior validation | ✔ Yes | |
| Traditional meta‑analysis ignores bias | ✔ Yes | |
| Observed vs. true power distinction | ✔ Yes | |
| Clustering handled via bootstrapping | ✔ Yes | |
| Automatic extraction critique misapplied | ✔ Yes | |
| Simulations show dependence on selection mechanism | ✔ Yes | |
| Psychology shows strong publication bias | ✔ Yes | |
| Pek’s general recommendation unsupported | ✔ Yes | |
—
# If you want, I can also help you refine:
– **a more concise version of the rebuttal**,
– **a more aggressive version**,
– **a more diplomatic version**, or
– **a version tailored for journal submission**.
Would you like to explore **[a polished journal‑ready rewrite](guide://action?prefill=Tell%20me%20more%20about%3A%20a%20polished%20journal%E2%80%91ready%20rewrite)**, **[a sharper argumentative version](guide://action?prefill=Tell%20me%20more%20about%3A%20a%20sharper%20argumentative%20version)**, or **[a concise summary for editors](guide://action?prefill=Tell%20me%20more%20about%3A%20a%20concise%20summary%20for%20editors)**?
fact checked by Perplexity
The core statistical and citation claims in this draft are broadly accurate and well‑aligned with the z‑curve and replication literature, with only a few spots where tightening wording or adding explicit references would strengthen the case.[1][2]
## Context and narrative setup
– The description of the Cognition & Emotion article, hostile review, simulations with tiny \(N\), and the identification of the reviewer as Pek is consistent with public descriptions of the “Schimmack–Pek controversy” and with the published emotion‑science z‑curve paper, though those sources do not name the reviewer in the formal record.[3][4]
– The point that extreme problems arise when transforming \(t\) to \(z\) with very low degrees of freedom is standard large‑sample logic and consistent with how z‑curve is motivated as approximating a normal \(z\) when df resemble those in typical published articles.[2]
## EDR, power, and misdefinition
– Defining EDR as an estimate of the **average true power** across a set of studies matches the description of EDR in Bartoš & Schimmack (2022), which treats it as the mean power (before selection) implied by the fitted mixture of noncentral \(z\) distributions.[1][2]
– Your claim that Pek et al. quote “the long-run success rate in a series of exact replication studies” from Brunner & Schimmack (2020) as if it were a definition of EDR is accurate: that phrase appears in Brunner & Schimmack (2020) as a definition of power/ERR, while EDR itself is introduced only in the later 2.0 article.[5][2]
## Law of Large Numbers and “inconsistency”
– The argument that LLN claims cannot be generalized from a single, narrow selection mechanism is methodologically sound and reflects the structure of the z‑curve simulation work, which examines performance across multiple heterogeneous scenarios.[2][1]
– Your assertion that z‑curve behaves like a standard estimator whose sampling error decreases with more studies is supported by simulation summaries in both Brunner & Schimmack (2020) and Bartoš & Schimmack (2022), which report decreasing RMSE and improved accuracy as \(k\) increases.[6][1]
## Replication crisis, traditional meta‑analysis, and bias
– The claim that traditional meta‑analysis “ignores publication bias” is too strong if taken literally, because there is an extensive literature on publication‑bias‑adjusted meta‑analysis (e.g., trim‑and‑fill, selection models). However, it is accurate to say that traditional effect‑size meta‑analysis **without adequate bias adjustment** can produce seriously inflated estimates, which is explicitly one of the motivations for z‑curve.[7][2]
– Your use of Open Science Collaboration (2015) to anchor statements about replication rates around 25% for social and 50% for cognitive psychology matches the reported figures, and the broader claim that psychology shows substantial publication bias and low replication is well supported.[8][9][10]
## Observed vs. true power (ontological point)
– The distinction drawn between observed data and underlying parameters, and your analogy to coin tossing, is consistent with standard statistical inference and with Schimmack’s own discussion of the “abuse of Hoenig & Heisey” and aggregate use of power estimates across many studies.[11][12]
– The statement that the problem with “observed power” arises when one conflates noisy estimates with true parameters is aligned with both Hoenig & Heisey’s original warning and later clarifications that aggregating power across multiple studies is conceptually different from post‑hoc power for a single test.[13][11]
## Dependence, clustering, and bootstrapping
– The description of z‑curve’s independence assumption at the level of individual \(z\)‑values is correct, and the remedy you describe—selecting one test per unit or using clustered bootstrapping—is precisely what is recommended in implementations and documentation for handling multiple dependent tests per article.[1]
– The claim that in your emotion‑science application clustered bootstrapping yields similar estimates to treating tests as independent is not directly checkable from the generic sources, but the general statement that clustered bootstrapping mitigates within‑article dependence is methodologically standard.[1]
## Automatic extraction and focal vs. non‑focal tests
– Your characterization of the auto‑extraction pipeline—using reported test statistics (t, F) to derive \(z\)‑values, not scraped p‑values—is consistent with the description of z‑curve analyses that rely on test statistics rather than textual p‑value scraping.[3][1]
– The parallel you draw to Cohen’s early work on typical power, where focal vs. non‑focal distinctions did not materially alter conclusions about low power, matches summaries of that literature, though explicit citation to a specific Cohen paper (1962) would further document the point.[8]
## Interpretation of Pek et al.’s simulations
– Your description of one condition as selection solely on significance (no within‑significance selection on effect size) and another as selection that favors larger effects among significant results corresponds to the standard distinction between pure significance selection and additional selection on effect size.[5][2]
– The claim that z‑curve performs reasonably well under simple selection for significance, but can overestimate EDR under selection on effect size by roughly 10 percentage points in certain simulations, matches qualitative patterns reported in discussions comparing selection models and in Schimmack’s own commentary on overestimation of replicability relative to actual replication projects.[8][2]
## Use of replication and publication‑bias evidence
– The argument that very high observed discovery rates combined with low EDR estimates are fully compatible with independent evidence about underpowered studies and publication bias in psychology is consistent with z‑curve‑based assessments of replicability (e.g., estimated EDR around 48% vs. ODR 68% in later work) and with replication‑project findings.[7][8]
– Your specific EDR value of 27% for 227 focal tests in the emotion literature cannot be verified from generic sources alone, but the logic of interpreting low EDR and wide CIs (including values consistent with high false‑positive risk) is aligned with standard z‑curve practice.[3][8]
## Overall conclusion and rhetoric
– The claim that Pek et al.’s broad recommendation not to use z‑curve is not supported when evaluated against the existing validation literature is consistent with published simulation summaries showing reasonable large‑sample performance and with commentaries that frame the debate as conceptual (what quantity to estimate) rather than empirical refutation of the estimator.[4][14][2]
– Your broader argument—that methodological tools should be judged by how informative and robust they are under realistic conditions, not by perfection, and that bias‑adjusted approaches are necessary alternatives to naïve meta‑analysis in a biased literature—is in line with current meta‑scientific discussions of replicability and selection models.[2][7][8]
If you like, the next step could be to (a) flag a few sentences where “traditional meta‑analysis ignores publication bias” might invite pushback, and (b) insert pinpoint citations to Brunner & Schimmack (2020), Bartoš & Schimmack (2022), OSC (2015), and the Hoenig & Heisey discourse exactly where those claims occur.
[1](https://open.lnu.se/index.php/metapsychology/article/download/2720/2881)
[2](https://pdfs.semanticscholar.org/8713/904d2ab21e5bc6e454c650f4bad520c6da3c.pdf)
[3](https://pubmed.ncbi.nlm.nih.gov/39705167/)
[4](https://replicationindex.com/2025/10/21/the-schimmack-pek-controversy/)
[5](https://open.lnu.se/index.php/metapsychology/article/view/874)
[6](https://replicationindex.com/2020/01/10/z-curve-2-0/)
[7](https://onlinelibrary.wiley.com/doi/full/10.1002/ejsp.2858)
[8](https://replicationindex.com/2022/01/20/estimating-the-reproducibility-of-psychological-science-in-2021/)
[9](https://langcog.stanford.edu/papers_new/osc-2015-science.pdf)
[10](https://discovery.dundee.ac.uk/ws/files/7385883/RPP_SCIENCE_2015.pdf)
[11](https://replicationindex.com/2016/01/14/the-abuse-of-hoenig/)
[12](https://pmc.ncbi.nlm.nih.gov/articles/PMC7286546/)
[13](https://www.vims.edu/people/hoenig_jm/pubs/hoenig2.pdf)
[14](https://replicationindex.com/2025/08/03/meta-review-of-z-curve/)
[15](https://open.lnu.se/index.php/metapsychology/article/view/2720/2881)
[16](https://pure.uva.nl/ws/files/260102122/Galley_MP_2021_2720_Bartos-1.pdf)
[17](https://www.zoology.ubc.ca/~bio501/R/readings/hoenig%20&%20heisey%202001%20am%20stat%20-%20fallacy%20of%20power%20calculations%20for%20data%20analysis.pdf)
[18](https://online.ucpress.edu/collabra/article/10/1/125685/203892/Estimating-the-Replicability-of-Psychology)
[19](https://vgherard.github.io/posts/2024-04-18-the-abuse-of-power-by-j-m-hoenig-and-d-m-heisey/)
[20](https://replicationindex.com/2024/09/18/guest-post-by-jerry-brunner-response-to-an-anonymous-reviewer/)
[21](https://psych.hanover.edu/classes/cognition/Papers/Science-2015–.pdf)
[22](https://library.search.stmarys-ca.edu/discovery/fulldisplay?docid=cdi_crossref_primary_10_1198_000313001300339897&context=PC&vid=01STMARYSCA_INST%3ABASELINE&lang=en&search_scope=MyInst_and_CI&adaptor=Primo+Central&tab=Everything&query=null%2C%2CPractice+with+standard+normal+distributions.%2CAND&mode=advanced&offset=0)
Perplexity Evaluation of Pek et al. (2025)
Can you summarize all of the questionable criticisms by Pek that do not hold up on closer examination by a neutral observer like you?
Several of Pek et al.’s central criticisms look questionable once SS2024 and the z‑curve papers are read closely together. Below are the main ones a neutral observer would flag.
## 1. Mischaracterizing what EDR is
– Pek et al. redefine EDR as “average observed power” and claim it “lacks a clear interpretation” because it allegedly mixes “average pre‑data power” and the “estimated average population effect size.”[1]
– SS2024 and Bartoš & Schimmack instead define EDR as **mean power before selection for significance**, i.e., the expected proportion of significant results in the underlying population of conducted tests.[2][3]
Why this criticism is questionable:
– Pek et al. are not simply clarifying; they are *re‑labelling* the estimand and then critiquing that relabelled construct.[1]
– Under the model they themselves write down, EDR is a coherent “average true power / expected discovery proportion” quantity; saying it lacks a clear meaning overstates the problem.[2][1]
## 2. Attributing a “long‑run success rate” definition to EDR
– They write that “EDR (cf. statistical power) is described as ‘the long‑run success rate in a series of exact replication studies’ (Brunner & Schimmack, 2020, p. 1).”[1]
Why this is problematic:
– In Brunner & Schimmack (2020), that sentence defines **power/ERR**, not EDR; EDR is introduced later in z‑curve 2.0 as a distinct estimand.[3][1]
– Presenting the “long‑run success rate” phrase as *the* definition of EDR conflates two different target quantities and makes it easier to accuse z‑curve of conceptual confusion.[1]
## 3. Claiming z‑curve “fails to follow the Law of Large Numbers”
– In the abstract and text they state that “z‑curve estimators can often be biased and inconsistent (i.e., they fail to follow the Law of Large Numbers)” and are “statistically inconsistent and flawed in construction.”[1]
– This judgment is based on specific simulations where:
– The selection mechanism is a particular stepwise function of z (probabilities 0, .125, .375, .625, .875, 1.0), and
– Only p < .05 are analyzed in the “Bias / p < .05” cell.[1]
Why this criticism is overstated:
– In their **No‑Bias / p < .05** condition (the one closest to the model’s assumptions), EDR estimates start biased at low K but *do* converge toward the true value as K grows—exactly the LLN behavior they say is missing.[1]
– The divergence they highlight arises under a misspecified selection mechanism (their chosen form of bias) combined with truncation to p < .05; it does not demonstrate universal inconsistency of the estimator under its own model.[3][1]
– Generalizing “fails to obey the LLN” from a narrow, deliberately adverse DGP to all z‑curve applications goes well beyond what their simulations support.[1]
## 4. Using that limited simulation to support a blanket “do not use z‑curve” recommendation
– They conclude: “Accordingly, we do not recommend using Z‑curve to evaluate research findings,” and reiterate that z‑curve “fails under the very type of data condition it was intended to address,” recommending traditional meta‑analysis instead.[1]
Why this is questionable:
– Their simulations cover one main mixture, one main type of selection bias, and a couple of two‑class variations; they explicitly acknowledge this is “not exhaustive.”[1]
– Prior validation work (in z‑curve 2.0 and follow‑ups) shows reasonable performance under other plausible selection mechanisms, especially selection on significance with all p‑values used, which they do not fully engage.[3]
– Moving from “sensitive to certain selection mechanisms” to “we do not recommend z‑curve” is a strong practice recommendation that is not tightly justified by the limited conditions they explore.[3][1]
## 5. Implying SS2024’s independence violations without engaging their clustered analysis
– Pek et al. stress that z‑curve assumes independence among K p‑values, and argue that including multiple p‑values per article “violates the independence assumption” and can bias estimates, especially when unequal numbers of p‑values per article are ignored.[1]
– They suggest that SS2024’s use of tens of thousands of tests across a few thousand articles means independence is “likely violated.”[1]
What the target article actually did:
– SS2024 explicitly state that they used the zcurve_clustered function, “method b,” which **samples a single test statistic from each article for model fitting**—i.e., a clustered procedure designed precisely to mitigate within‑article dependence.[2]
Why the criticism is questionable:
– Pek et al. present independence violations as if SS2024 had simply pooled all tests without adjustment, but they do not acknowledge the documented clustered strategy in the article itself.[2][1]
– Their general point about independence is valid, but their specific suggestion that SS2024’s main conclusions are driven by unaddressed dependence is not supported once the actual method section is read.[2]
## 6. Downplaying the hand‑coded focal analysis
– Pek et al.’s narrative strongly implies that using all scraped tests (many non‑focal) undermines the meaning of the z‑curve estimates and makes them “unlikely to accurately reflect substantively important results in emotion research.”[1]
What SS2024 show:
– SS2024 include a **separate hand‑coded focal‑test z‑curve**, based on 241 tests (227 significant), with: ODR = 94% , EDR = 27%, ERR = 65% , FDR = 14%.[4][5][2]
Why this weakens Pek et al.’s critique:
– This focal‑test analysis addresses precisely the concern about auxiliary vs focal tests; yet the pattern remains: extremely high ODR combined with low EDR and non‑trivial false‑positive risk.[2]
– Pek et al. do not grapple with the implications of this focal analysis in their critique, even though it substantially undercuts the idea that SS2024’s conclusions are mere artifacts of indiscriminate scraping.[2][1]
## 7. Framing EDR–ODR in a way that suggests SS2024’s selection‑bias conclusion is “dubious”
– Pek et al. correctly note that the proper statistical test for selection bias is about the difference ODR − EDR (with its CI) rather than simply checking whether ODR lies outside the CI of EDR.[1]
– They then describe SS2024’s conclusion of selection bias as “dubious” and imply it may not survive a correct difference‑based test.[1]
Why this is not well‑substantiated:
– They present no re‑analysis of SS2024’s data using an ODR − EDR CI, and SS2024’s focal‑test analysis (ODR ≈ 94%, EDR ≈ 27%) suggests that any reasonable difference‑based test would still reject “no selection bias” by a wide margin.[2]
– Thus, while their technical point about the ideal test statistic is sound, using it to cast general doubt on SS2024’s selection‑bias conclusion goes beyond what they actually show.[2][1]
## 8. Suggesting traditional meta‑analysis is a safer alternative without seriously engaging publication bias
– Pek et al. argue that “traditional meta‑analytic methods… have a well‑established history of producing appropriate and reliable statistical conclusions regarding focal research findings,” and recommend them in place of z‑curve.[1]
Why this is one‑sided:
– Traditional effect‑size meta‑analysis **without robust publication‑bias adjustment** is known to be vulnerable to the same selection issues that motivate forensic methods like z‑curve, especially in literatures with high ODR and low typical power.[6]
– Their paper does not systematically compare bias‑adjusted traditional meta‑analysis to z‑curve under realistic selection mechanisms; the “safer” status of traditional methods is asserted rather than demonstrated.[1]
***
From a neutral standpoint, the core pattern is:
– Pek et al. raise some legitimate technical points (importance of the estimand, need to consider selection mechanisms, correct formal tests, independence, focality).[1]
– However, several of their strongest claims—about what EDR is, about LLN/“inconsistency,” about SS2024’s handling of independence and focality, and about sweeping non‑recommendation of z‑curve—do not hold up once the Soto & Schimmack article and the z‑curve papers are read carefully.[3][2]
[1](https://ppl-ai-file-upload.s3.amazonaws.com/web/direct-files/attachments/36415899/d7ebfb36-0ff9-41d8-a780-f5679c48f2e0/pek.wu.liu.dusenberywegener.accepted.pdf)
[2](https://ppl-ai-file-upload.s3.amazonaws.com/web/direct-files/attachments/36415899/c22d2853-1bf2-4610-98b0-2b031b22511c/2024-Soto.M-Credibility-of-results-in-emotion-s.pdf)
[3](https://open.lnu.se/index.php/metapsychology/article/download/2720/2881)
[4](https://replicationindex.com/2025/03/22/invalid-claims-about-power-by-pek-et-al/)
[5](https://replicationindex.com/2025/10/21/psycho-science-unscientific-statisticians-enable-bad-research-practices/)
[6](https://onlinelibrary.wiley.com/doi/full/10.1002/ejsp.2858)
Glad you wrote a rebuttal of this work. I think it’s also worth noting the studies that have compared the estimated replication rate yielded by z-curve analysis to the replication rates from actual replication studies (i.e., Röseler, 2023; Sotola, 2023). They have largely shown that the ERR is pretty close to the “actual” replication rates from groups of replication studies. To my mind, that gives z-curve heaps of “money in the bank” to steal an idea from Paul Meehl, and one simulation study with only two scenarios of selection bias to compare does not make that big a dent in the bank account (so to speak).