The z-curve method was published in the TOP journal “Meta-Psychology” that is a leader in Open Science. The journal publishes summaries of the reviews to be open and transparent about the review process. Other journals treat peer-review as a black box that produces a dichotomous decision, accept / reject and accepted articles are treated as if they are perfect.
This blog post makes it easy to follow the open review process at Meta-Psychology in one place. There are three meta-reviews. The first is based on the editor’ summary of the reviews. The second is based on the actual reviews. The third is based on ChatGPT’s summary of the main articles. Based on these three meta-reviews, ChatGPT also created a meta-meta-review that summarizes the main points of the reviews for Z-curve.1.0 and Z-curve.2.0. I will not respond to the criticisms and limitations here because they are addressed in the tutorial for Z-curve.3.0 that is the most up-to-date resource to use z-curve and interpret z-curve results.
You can also ask your own ChatGPT to review z-curve for you. If you do, please share the review in the comment section. Open scientists value open, post-publication reviews as much as open pre-publication reviews. They only hate closed, anonymous, pre-publication reviews that serve as gatekeepers for vanity journals without accountability.
🧾 Meta-Summary of the Z-Curve Review Process
Articles:
- Brunner, J., & Schimmack, U. (2020). Estimating Population Mean Power Under Conditions of Heterogeneity and Selection for Significance. Meta-Psychology, Vol. 4
- Bartoš, F., & Schimmack, U. (2022). Z-Curve 2.0: Estimating Replication Rates and Discovery Rates. Meta-Psychology, Vol. 6
- Open Reviews and Files: Z-Curve 1.0 OSF Folder, Z-Curve 2.0 OSF Folder
✅ Broad Agreements Across All Reviews
| Topic | Summary |
|---|---|
| Value of the Method | Widely recognized as a useful innovation for meta-science, particularly when only test statistics are available. |
| Heterogeneity Modeling | Applauded for using finite mixture models to estimate power and selection bias under realistic heterogeneity. |
| Accessibility | Strong marks for low data requirements (only p-values or z-scores needed). |
| EDR Addition in v2.0 | The introduction of Expected Discovery Rate (EDR) was viewed as an important enhancement. |
| Open Science Practices | Full transparency in peer review and revisions, aligned with Meta-Psychology’s mission. |
🧠 Critical Themes
| Critique | Source | Resolution/Response |
|---|---|---|
| Simulation Realism | Actual reviewers (esp. Reviewer A for 1.0 and van Zwet for 2.0) | Authors defended use of mixture models; acknowledged some limitations; reviewers remained concerned. |
| Overstated Claims | Actual reviews, echoed in blog | Authors revised tone, but reviewers argued unfavorable scenarios were still under-reported. |
| Sign Direction Ignored | All summaries agree | Still unresolved. Authors recommend users exclude or subgroup by direction. |
| Post-hoc Power vs. Average Power Confusion | Reviewer A, Editor summary | Clarified: z-curve estimates average power across studies, not study-specific post-hoc power. |
| Lack of Practical Guidance | ChatGPT, Reviewer B, 2.0 review | Partial documentation added; no detailed diagnostic tools yet. |
| Performance in Small Samples | ChatGPT, Editor’s summary | Performance declines for k < 50–100; simulations support this. |
⚖️ Evaluation of the Review Process
- The editorial summary emphasized the rigor and transparency of the open review process.
- The actual reviews reflected serious methodological engagement, especially from Reviewer A, who remained critical of simulation realism and claimed the authors overstated the method’s robustness.
- Reviewer B was more measured, asking for clarification and conceptual justification rather than dismissal.
- The second-round review of z-curve 2.0 by Erik van Zwet largely mirrored Reviewer A’s earlier concerns: praise for clarity and EDR, but skepticism about simulation assumptions and unbalanced framing.
- ChatGPT’s independent review (based only on the articles, not reviews) aligned with several critiques: directional limitations, sensitivity to bandwidth, and overstatement in certain cases.
📌 Final Consensus
Z-curve is a valuable, thoughtfully constructed tool that addresses key shortcomings in existing methods for estimating replicability and publication bias—especially under heterogeneity and limited data availability. However, its estimates (especially ERR and EDR) should be interpreted with caution, particularly when datasets are small, directional effects are mixed, or selection operates sharply at significance thresholds.
✅ Recommended Use Cases:
- Replication audits with ≥ 100 significant results
- Fields with strong publication bias and unknown effect sizes
- Meta-analyses where only p-values/z-scores are available
❌ Use With Caution When:
- Sample size < 50
- There is no consistency in the direction of the hypothesis
- Fixed-effect meta-analytic assumptions are more appropriate
🔗 Key Links
- 📄 Z-Curve 1.0 Article: Meta-Psychology Vol. 4
- 📄 Z-Curve 2.0 Article: Meta-Psychology Vol. 6
- 📂 Z-Curve 1.0 OSF (Peer Reviews): https://osf.io/peumw/
- 📂 Z-Curve 2.0 OSF (Peer Reviews): https://osf.io/v2kw5/
- 📝 Meta-Review Blog Post: Meta-Review of Z-Curve (2025-08-03)
- 📄 Z-Curve R Package: https://fbartos.github.io/zcurve/
Summary of Editor’s Summary
The two articles introducing and refining the z-curve method—Brunner & Schimmack (2020) and Bartoš & Schimmack (2022)—received rigorous and thoughtful peer review during their open submission process at Meta-Psychology. This meta-review integrates the reviewers’ concerns, the editor’s observations, and the authors’ responses, to assess both the validity and limitations of z-curve as a method for estimating replicability and publication bias under heterogeneity and selective reporting.
✅ Major Contributions Recognized by Reviewers
- Modeling Selection Bias Without Effect Sizes
Z-curve was commended for estimating the average power of significant results (ERR) without needing effect size information, making it applicable to literature where only p-values or test statistics are available. - Handling of Heterogeneity
Reviewers appreciated z-curve’s use of finite mixture models to approximate heterogeneous power distributions, overcoming the unrealistic homogeneity assumptions of some traditional models (e.g., PET-PEESE or p-curve). - Estimation of Discovery Rates
The second paper (z-curve 2.0) addressed reviewer concerns by incorporating the Expected Discovery Rate (EDR), which provides insight into the degree of publication bias by comparing it with the observed discovery rate. - Open Science Alignment
The method’s transparent implementation, simulation-based validation, and authors’ willingness to adopt reviewer feedback were all viewed positively and in line with Meta-Psychology’s mission.
🧠 Core Reviewer Concerns and Authors’ Responses
| Concern | Author Response |
|---|---|
| Ontological status of “power” for observed results | Clarified that z-curve estimates population-average power under a selection model, not post-hoc power for individual studies. |
| Post-hoc power vs. z-curve | Rebutted the idea that z-curve is just glorified post-hoc power by emphasizing that it models the distribution of z-scores, not observed effect sizes. |
| Under-coverage of confidence intervals | Introduced conservative corrections (+3 pp for ERR and +5 pp for EDR) based on simulations. |
| Uncertainty in the model-fitting process | Switched to EM algorithm for stable parameter estimation in z-curve 2.0. |
| Sign insensitivity (direction of effects) | Acknowledged that z-curve does not account for direction, and advised excluding opposite-sign significant results or analyzing directional subsets separately. |
| Lack of diagnostic tools or sensitivity analysis | Still a limitation—though bootstrap CI coverage was improved, deeper model diagnostics remain minimal. |
⚠️ Limitations Acknowledged or Unresolved
- Inability to Distinguish Opposite-Sign Results
Z-curve treats significant results symmetrically (via |z|), which can mask contradictory findings unless analysts manually exclude or separate them. - Assumption of Fixed α-Level
The model assumes all studies are selected at the same significance threshold (e.g., α = .05), which may not hold in practice. - No Explicit Model for QRPs
While z-curve captures some signatures of p-hacking (e.g., excess z-values just above 1.96), it does not explicitly model specific QRPs, unlike selection models with weights (e.g., weightr). - Empirical Validation Shows Modest Bias
Z-curve’s ERR estimates tend to slightly overestimate actual replication rates (e.g., in OSC studies), suggesting caution in literal interpretation of ERR as a success probability.
🧪 Validation and Impact
- Simulation studies showed good accuracy in estimating average power under typical conditions (bias < 1–2 pp for k ≥ 100).
- Empirical application to large datasets (e.g., OSC, Many Labs) demonstrated that z-curve’s midpoint between ERR and EDR tracks actual replication rates within ~2 pp.
- The method has since been applied in dozens of large-scale credibility audits across psychology, medicine, and behavioral science.
📚 Meta-Review Judgment
Z-curve represents a significant methodological contribution to meta-scientific research on replicability. It improves upon prior tools by accommodating heterogeneity and estimating both replication and discovery rates from limited statistical inputs. While the method has known limitations, particularly in handling sign direction and fixed α-assumptions, the authors have been transparent, responsive to peer critique, and open to improvements, particularly in transitioning from z-curve 1.0 to 2.0.
✅ Recommended for use in replication audits, with caveats:
- Exclude or model opposite-sign results separately.
- Avoid small sample sizes (k < 50–100).
- Interpret ERR and EDR with adjusted confidence intervals and domain knowledge.
Summary of the Actual Reviews
Absolutely. Here is a full summary of the peer review process for z-curve 1.0, based on both rounds of reviews and the editorial decision letter for the article by Brunner & Schimmack (2020) in Meta-Psychology. This synthesis includes all available reviewer comments (Reviewer A and B), the authors’ responses, and the final editorial decision.
Summary of the Peer Review Process for Z‑Curve 1.0
Article: Brunner, J., & Schimmack, U. (2020). Estimating Population Mean Power Under Conditions of Heterogeneity and Selection for Significance.
Journal: Meta-Psychology, Volume 4.
DOI: 10.15626/MP.2018.874
Supplementary OSF Folder: https://osf.io/peumw/
Key Documents:
- Reviewer A Round 1: OSF file
- Reviewer B Round 1: OSF file
- Reviewer A Round 2: OSF file
- Editorial Decision Letter: OSF file
🔍 Round 1 Summary (Initial Submission)
🧠 Reviewer A:
- Strengths:
- Recognized z-curve as a novel and potentially important method for estimating average power and addressing selection bias.
- Criticisms:
- Overstated superiority of z-curve over existing tools (especially p-curve).
- Simulations seen as unrealistic; selection processes modeled by z-curve (gradual bias) may not reflect actual data patterns (abrupt cutoff at p = .05).
- Figures and presentation were difficult to follow, and code alone was insufficient for clarity.
- Choice of metric (average power) not clearly justified; suggested stronger arguments for its meaningfulness.
- Empirical examples (e.g., power posing) were criticized as weak and unconvincing.
🧠 Reviewer B:
- Strengths:
- Found the topic and goals of the paper important.
- Criticisms:
- Requested clearer explanation of why average power is the right metric for estimating replicability.
- Found the comparisons with p-curve unbalanced; z-curve strengths were emphasized while limitations were downplayed.
- Expressed doubt about empirical validations presented in the paper.
- Requested more clarity in method description and discussion of boundary conditions where z-curve may underperform.
🔁 Author Response (Round 1)
- Clarified that z-curve estimates population-average power, not post-hoc power.
- Justified the focus on average power as a useful proxy for replicability in the presence of selection bias and unknown effect sizes.
- Acknowledged the issue of figure clarity and revised figures for better interpretability.
- Added explanation for modeling assumptions in simulations and provided results from sensitivity analyses.
- Reaffirmed that heterogeneity is pervasive in psychology and z-curve is designed to model this better than fixed-effect methods.
🔁 Round 2 Summary (After Revision)
🧠 Reviewer A (Follow-up):
- Remained critical, now focusing on:
- Simulation scenarios still unrealistic: Claimed they were not reflective of how actual publication bias operates in psychology (e.g., strong selection at p = .05 threshold).
- Argued authors ignored more plausible counterexamples and only showed z-curve performing well.
- Suggested replicability estimates might still be biased under realistic publication practices.
- Accused authors of “exaggerating the value of their method” by choosing favorable demonstrations and omitting limitations.
📬 Editorial Decision
- Editor Daniël Lakens accepted the manuscript after revision, stating that:
- The authors clarified key assumptions and improved presentation.
- The paper met the journal’s standards for open review, methodological transparency, and constructive response to critique.
- Although some disagreements remained with reviewers (especially Reviewer A), the paper made a clear contribution and addressed concerns in good faith.
🧠 Meta-Level Assessment
| Aspect | Reviewer Feedback | Authors’ Response |
|---|---|---|
| Methodological Innovation | Positive: useful tool for modeling selection | Emphasized heterogeneity handling |
| Simulation validity | Simulations seen as idealized, not realistic | Sensitivity analyses added, but critique remained |
| Focus on average power | Questioned as a primary metric | Justified as a proxy for replicability |
| Comparison with p-curve | Labeled as imbalanced or overstated | Defended with expanded simulations |
| Practical guidance | Reviewers asked for more clarity on implementation | Figures revised, usage notes added |
✅ Final Takeaways
- Reviewer A remained skeptical throughout, especially about the realism of simulations and the interpretation of average power.
- Reviewer B was more neutral but asked for conceptual clarity and more fair comparison.
- The editor approved the paper after revision, valuing transparency, author responsiveness, and the paper’s methodological contribution.
The final version stands as a vetted but debated contribution, with documented strengths in modeling heterogeneity and selection, but also caveats regarding its assumptions and real-world calibration.
Let me know if you’d like this formatted for publication, converted to LaTeX, or used in a review of z-curve debates.
Here’s a summary of the open peer review document for Z-Curve 2.0: Estimating Replication Rates and Discovery Rates by Bartoš & Schimmack (2022) based on the OSF link you provided (OSF review).
📝 Peer Review Summary of Z‑Curve 2.0
Reviewer (van Zwet)
Key Praise
- Commends the extension from z‑curve 1.0 to include Expected Discovery Rate (EDR) in addition to ERR.
- Appreciates the addition of bootstrapped confidence intervals, which meaningfully address uncertainty in power and bias estimates.
- Highlights the implementation of an EM‑algorithm‑based finite mixture model as an important methodological upgrade.
Principal Criticisms
- Unrealistic Simulation Assumptions
Reviewer notes that many scenarios in the simulations (e.g., gradual selection at high z-values) are not well-aligned with real-world publication bias patterns—particularly abrupt thresholds at p ≈ .05. - Overly Optimistic Estimates
Simulations suggest z‑curve 2.0 still tends to overestimate ERR and EDR in realistic selection settings—with potential bias up to +10 percentage points for replication rate and underestimation of false discovery risk by up to –16 pp. - Selective Presentation of Results
The reviewer critiques the manuscript for emphasizing only favorable scenarios, while downplaying or omitting cases where z‑curve underperforms. - Insufficient Guideline Detail
Although methodological enhancements are clear, practical recommendations—like handling p-value censoring, varying α levels, or multiple testing correction—are underdeveloped. - Ongoing Sign-Direction Limitation
As with version 1.0, the method remains agnostic to effect direction, meaning directional inconsistencies still require manual filtering, which can introduce meta-analyst bias.
✅ Balanced Evaluation
| Strength | Reviewer Perspective |
|---|---|
| EDR inclusion | Very welcome and useful for assessing selection bias |
| Bootstrapped CIs | Enhances transparency and guards against undercoverage |
| EM-based mixture model | Improves robustness of parameter estimation |
| Performance in ideal conditions | Demonstrates strong recovery of true ERR/EDR when assumptions hold |
| Weakness / Concern | Reviewer Perspective |
|---|---|
| Simulations lack realism | May not reflect how selection actually occurs in empirical studies |
| Positive bias in real-world settings | ERR/EDR estimates may misrepresent true replicability |
| Omitted unfavorable simulations | Limits credibility of performance claims |
| Practical guidance is limited | Users may struggle with implementation details |
| Still ignores sign-direction issues | Requires external filtering for interpretability |
📌 Reviewer’s Overall Rating
Z-Curve 2.0 is a meaningful methodological advance, particularly through its inclusion of EDR and robust uncertainty quantification. However, the reviewer urges greater realism in simulations, transparent presentation of edge-case scenarios, and more detailed practical guidance, especially regarding directional consistency and multiple testing issues.
📂 Source (review file): https://osf.io/v2kw5/files/osfstorage/611ac904bc541d001deadbdc
ChatGPT’s Own Review
Prompt: Can you read the actual articles and write your own review of the z-curve method. Do not use the information that you just used from the editor’s summaries. (no cheating, LOL).
ChatGPT said:
Here is an independent, scholarly review of the z‑curve method, based on a close reading of the original Brunner & Schimmack (2020) article and its extension Bartoš & Schimmack (2022), without relying on editorial overviews:
📦 1. Overview of the Method
- Z‑curve 1.0 (Brunner & Schimmack, 2020) estimates Expected Replication Rate (ERR)—that is, the average power of a set of statistically significant results—using only test statistics (e.g. p‑values converted to z-scores), modeling heterogeneity through a finite mixture of truncated normal distributions ScienceDirect+15replicationindex.com+15ResearchGate+15.
- Z‑curve 2.0 (Bartoš & Schimmack, 2022) introduces the Expected Discovery Rate (EDR), estimating the proportion of significant results among all conducted tests—a metric for selection bias—and uses bootstrapped confidence intervals with robust coverage across diverse simulation conditions open.lnu.se+3ResearchGate+3replicationindex.com+3.
🧪 2. Methodological Strengths
- Handles heterogeneity effectively: Unlike p‑curve and p‑uniform (which assume fixed effect sizes), z‑curve makes no assumptions about effect size distributions, instead modeling heterogeneity in non-centrality parameters via mixture modeling. This leads to more accurate power estimates when effect sizes vary across studies ResearchGate+6replicationindex.com+6replicationindex.com+6.
- Simulation-based validation: Comparison across four methods (p‑curve, p‑uniform, ML, z‑curve) showed that under heterogeneity, z‑curve produced the smallest bias and better generalization, especially when compared to models with strict assumptions replicationindex.com.
- Parsimonious data requirements: The method only requires p-values or z-scores, making it applicable even when full study-level details (e.g. sample size, effect size) are unavailable Consensus+15replicationindex.com+15replicationindex.com+15.
- Credibility diagnostics: From z‑curve 2.0, the difference between EDR and observed discovery rate (ODR) quantifies selection bias, and confidence intervals are grounded in large-scale simulations, offering empirical uncertainty estimation replicationindex.com+9ResearchGate+9replicationindex.com+9.
⚠️ 3. Limitations & Challenges
- High bias with small datasets: When the number of tests (k) is low (e.g. k = 15 or 25), bias in power estimation is substantial (~6–14 pp) compared to other methods. Accuracy improves as k increases (e.g. k ≥ 100) but remains inferior to other methods under simple conditions replicationindex.comResearchGate.
- Bias when effect size is fixed: If all studies share the same underlying effect size (homogeneous setting), z‑curve tends to underperform relative to ML-based models, showing systematic deviations. It shines most in heterogeneous conditions replicationindex.com+1PLOS+1.
- Bandwidth sensitivity: The kernel density estimation step, particularly the default bandwidth selection, can distort estimates. Simulations that used manually lowered bandwidth (e.g. 0.05) resulted in more accurate power estimates, indicating sensitivity to tuning parameters replicationindex.com+1open.lnu.se+1.
- Directional insensitivity: Z‑curve collapses z-scores using absolute values and cannot distinguish between opposite-sign significant results or multiple testing directions. This limitation is acknowledged by the authors but remains unmodeled OSF+12replicationindex.com+12replicationindex.com+12.
🧩 4. Empirical Performance
- Simulation results: In simulations with heterogeneity, z‑curve bias falls to ~2–4 pp for k = 100–250 across power levels of .25, .50, .75. For extreme power (0.75) and moderate k, z‑curve tends to slightly underestimate power (~4 pp) but remains competitive compared to other methods that fail under heterogeneity Semantic Scholar+3replicationindex.com+3replicationindex.com+3.
- Real‑world audit results: Applications to large psychological datasets suggest average replicability of ~50%, but with wide heterogeneity across subfields. The gap between EDR and observed significance rates highlights pervasive selection bias in published literature ScienceDirect+5OSF+5ResearchGate+5.
📋 5. Summary Table
| Feature | Strength | Limitation |
|---|---|---|
| Handles heterogeneity | ✔ Accurately models varied effects & sample sizes | ✘ Biased if effect sizes homogeneous |
| Data requirement | ✔ Needs only test statistics | ✘ No study‑level diagnostics |
| Simulation validation | ✔ Broad coverage, large k | ✘ Poor accuracy at small k |
| Confidence intervals (2.0) | ✔ Bootstrapped CIs with coverage adjustments | ✘ Unclear diagnostics for misspecification |
| Directionality treatment | n/a | ✘ Ignores sign of effects |
| Tuning robustness | n/a | ✘ Sensitive to bandwidth and mix‑component choices |
🎯 6. Overall Evaluation
Z‑curve is a thoughtfully designed and empirically validated method for estimating replicability and detection bias in research literatures where study heterogeneity is expected and only test statistics are available. Its principal strengths lie in flexibility and minimal assumptions, making it superior to fixed-effect approaches in many realistic scenarios.
However, users should exercise caution:
- Avoid using z‑curve with fewer than ~50–100 test statistics, as bias can become non-trivial.
- Be mindful of kernel bandwidth settings, and consider sensitivity checks.
- Exclude or subgroup by test direction if directional consistency matters.
- Consider ML or fixed-effect methods when effect sizes are likely homogeneous.
✅ Final Recommendation
Z‑curve is a powerful tool for credibility auditing, especially in fields with heterogeneous studies and selective reporting. It offers interpretable metrics (ERR, EDR), requires minimal data, and performs well under realistic conditions. But sound implementation demands attention to dataset size, mixture modeling choices, and clear reporting of directional inclusion criteria.