Pith. sign in

REVIEW 3 minor 33 references

Outlier-handling choices flip the verdict in 11.5% of behavioral science meta-analyses.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:22 UTC pith:NOR6YTPH

load-bearing objection Solid, pre-registered empirical benchmark: first large-scale comparison of outlier-handling treatments in meta-analysis; the one soft spot (baseline integrity) is disclosed and likely conservative.

arxiv 2607.23174 v1 pith:NOR6YTPH submitted 2026-07-25 econ.EM

Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses

classification econ.EM
keywords meta-analysisoutlier handlinginfluence diagnosticsresearcher degrees of freedompre-registrationsensitivity analysissmallest effect size of interestbehavioral science
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether a meta-analyst's decision about what to do with extreme or highly influential estimates changes the conclusions of research syntheses. Across 358 behavioral science meta-analyses, four pre-registered outlier treatments barely move the pooled effect size—the median absolute change in Cohen's d is at most 0.047—yet at least one treatment reverses the statistical-significance verdict in 11.5% of the meta-analyses and the smallest-effect-of-interest (|d| ≥ 0.20) verdict in 15.9%. The reversals concentrate almost entirely among results already near the decision boundary; strongly significant results essentially never change. The authors read this as evidence that outlier handling is a researcher degree of freedom, and argue that pre-registering the exact handling rule is a practical safeguard.

Core claim

The central claim is that in behavioral science meta-analyses with at least ten estimates, outlier and influence handling decisions—dropping the most extreme estimate, deleting estimates flagged by studentized deleted residuals or DFBETAS, or winsorizing the tails—have little effect on the magnitude of the pooled mean (median |Δd| ≤ 0.047), but a meaningful effect on categorical interpretation. Counting either of two estimators and any of the four treatments, 7.7% of the 715 estimable estimator-by-meta-analysis cells change statistical significance and 10.3% change smallest-effect-of-interest status; at the level of meta-analyses the rates are 11.5% and 15.9%. The authors emphasize that wins

What carries the argument

The comparison machinery is a pre-registered grid: each of 358 meta-analyses is estimated under two estimators—random-effects with a small-sample t correction, and unrestricted weighted least squares (an inverse-variance weighted mean with heterogeneity-inflated standard errors)—and each is compared against a 'do nothing' baseline on four treatments: dropping the single most extreme effect, deleting estimates with studentized deleted residuals above 3, winsorizing at the 5th and 95th percentiles, and deleting estimates with |DFBETAS| above 2/√k. Treating one handling choice at a time while holding thresholds and edge rules fixed makes any observed change attributable to the handling rule. Th

Load-bearing premise

The 'do nothing' baseline is assumed to reflect genuinely unhandled data; if the source meta-analyses had already removed or corrected extreme estimates, the measured reversal rates understate how much outlier handling matters.

What would settle it

One finding that would settle it: re-run the pre-registered treatments on a set of meta-analyses where raw primary-study data are available and the original authors' outlier removals are documented, and build the do-nothing baseline from the unhandled raw estimates. If reversal rates fall to near zero, the reported 11.5% and 15.9% figures are an artifact of baselines that already incorporate outlier handling. Conversely, if reversals were found concentrated among strongly significant or large effects—which the paper claims essentially never change—the central claim would also be contradicted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Meta-analysts should pre-specify their outlier handling rule: a defensible choice of method can change the headline verdict in about one in nine syntheses.
  • Reviewers should ask for sensitivity analyses that report the central result both with and without outlier treatment; borderline findings are where the risk of reversal lives.
  • Winsorizing is the least disruptive handling option and DFBETAS-based deletion the most; a meta-analyst wanting a conservative check can use winsorizing, while DFBETAS under unrestricted weighted least squares is the most sensitive diagnostic to report.
  • Because reversal rates rise when meta-analyses are small (k below 20) and fall when they are large, outlier handling deserves special attention in small syntheses.
  • These rates are a benchmark: a 7.7% cell-level and 11.5% meta-analysis-level significance-reversal rate is comparable in scale to the discordance previously documented for heterogeneity-estimator choice, so outlier handling is no smaller a degree of freedom.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The reversal rates are measured against a 'do nothing' baseline taken from compiled meta-analyses that may already have had outliers removed by their original authors; if so, the true sensitivity of meta-analytic conclusions to outlier handling is understated, not overstated.
  • Because the sample is limited to behavioral science syntheses with at least ten effects, applying the same pre-registered pipeline to medical or other disciplinary meta-analyses, where effect distributions and typical k differ, could meaningfully change the rates; that is a testable extension of the design.
  • The fact that winsorizing only ever moves results toward statistical significance suggests a caution: a method that shrinks tails without asking why an estimate is extreme could make a synthesis look more conclusive than the data warrant, especially in areas with publication bias.
  • A direct practical test of this design's value: for a set of meta-analyses with fully documented raw data, compare the conclusions reached by teams that pre-registered versus teams that did not, to see whether pre-specification actually reduces outcome-dependent outlier handling.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. This paper quantifies how pre-registered outlier and influence-handling decisions affect meta-analytic conclusions in 358 behavioral-science meta-analyses (k≥10; 16,664 estimates) from the BEAR v2 collection spanning psychology, psychotherapy, and exercise. Four active treatments—dropping the most extreme |d|, deleting estimates with studentized deleted residuals >3, winsorizing at the 5th/95th percentiles, and deleting estimates with |DFBETAS|>2/√k—are compared against a do-nothing baseline under two estimators: random effects (REML with Hartung–Knapp adjustment) and unrestricted weighted least squares (UWLS). Outcomes are absolute change in pooled d, statistical significance (p<.05), and whether |pooled d| reaches a SESOI of ≥0.20. The median absolute change in d is at most 0.047 across treatments. At least one treatment changes statistical significance in 7.7% of 715 estimable estimator-by-meta-analysis cells (11.5% of 358 meta-analyses) and changes SESOI status in 10.3% of cells (15.9% of meta-analyses). Reversals are largely confined to results near the decision boundary; DFBETAS is the most sensitive and winsorizing the least. Robustness checks (k≥20, k≥5, alternative cutoffs, one-estimate-per-study) are consistent. The entire pipeline was pre-registered, the dataset is frozen, and a replication package is provided.

Significance. If correct, this is the first large-scale empirical reference for the sensitivity of meta-analytic significance and substantive-size verdicts to outlier handling, an under-reported researcher degree of freedom. The paper's strengths are its design: a pre-analysis plan with fixed thresholds and edge rules, a frozen dataset with checksum, and a publicly archived replication package that makes the confirmatory results push-button reproducible. The comparison to Langan et al.'s heterogeneity-estimator discordance (7.5%) provides a useful benchmark. The descriptive nature is appropriate; the authors do not claim to identify which treatment is correct. Limitations—possible prior cleaning in the source databases, repeated d/SE rows of unknown provenance in psychology, and the conflation of small-study effects with true outliers—are acknowledged and partially mitigated by the retention of extreme values (|d|>10) in the psychology source. The findings should be of interest to applied meta-analysts, methods researchers, and journal reviewers.

minor comments (3)
  1. [§5.2, Limitations] The statement that prior cleaning in the source databases would make the treatment effects "understate how much these choices matter" is plausible but not logically guaranteed: if earlier cleaning selectively removed estimates that were propping up significance, the remaining baseline could be more fragile, and the reported reversal rates could overstate sensitivity. The retention of 37 estimates with |d|>10 supports the conservative-direction reading, but a sentence acknowledging the alternative direction and noting that a direct test would require access to source pre-processing would be more precise.
  2. [Table 4, notes] The column "Not estimable" in Table 4 is not defined in the table notes. The text in §4.4 explains that these arise from metafor REML non-convergence, but the table note should carry this definition for readers who start with the tables.
  3. [§3.3, Outcomes] The SESOI threshold c=0.20 is based on Cohen's small effect for d. Because the psychology studies are converted from Fisher-z correlations, it may be worth adding a sentence clarifying whether the threshold is intended to be directly comparable across the three fields, which have different original metrics.

Circularity Check

0 steps flagged

No circularity: the paper reports preregistered measurements on a fixed dataset, not predictions derived from a fitted model.

full rationale

The paper's central results are direct empirical summaries: it applies four fixed outlier-handling treatments and two estimators to a frozen dataset, counts how often the statistical-significance or SESOI status changes relative to a do-nothing baseline, and reports median absolute changes in Cohen's d. There is no estimation-then-prediction loop: every treatment, threshold, estimator, and outcome was fixed in the pre-analysis plan before outcomes were computed, so the reversal rates cannot be a product of post hoc fitting. The estimators (RE with HKSJ and UWLS) are defined explicitly in Section 3.2, and UWLS is justified by the Gauss-Markov theorem rather than by an imported uniqueness claim. Although some citations to the authors' own prior work motivate UWLS and supply one of the three source collections, these are not load-bearing in the sense of forcing the result; the exercise-data overlap is disclosed in the Competing Interest statement and affects sample composition, not the derivation of the measured statistics. The Discussion's statement that DFBETAS should have the most effect 'because this is what DFBETAS measures' is an acknowledged expectation from the definition of the diagnostic, not a predicted result that is then treated as independently derived. The Section 5.2 limitation that the 'do-nothing' baseline may already contain prior outlier-handling decisions is a measurement/validity concern about the interpretation of the reversal rates; it does not make any reported rate equivalent to an input by construction. Overall, the paper is a self-contained empirical sensitivity survey with preregistered comparisons, and no circular step is exhibited.

Axiom & Free-Parameter Ledger

5 free parameters · 3 axioms · 0 invented entities

The reversal rates are direct measurements, not fitted predictions. The free parameters are pre-registered design choices, not fit to data; the axioms are the representativeness of the BEAR collection, the integrity of the baseline, and the appropriateness of the two estimators. No new theoretical entities are introduced.

free parameters (5)
  • Studentized deleted residual cutoff = 3
    Pre-registered threshold for deleting estimates; intentionally stricter than Viechtbauer & Cheung's illustrative 1.96. Robustness checked at 2.5.
  • DFBETAS deletion threshold = 2/sqrt(k)
    Pre-registered Belsley-Kuh-Welsch-style cutoff for influential estimates. Robustness checked at 3/sqrt(k).
  • Winsorizing percentiles = 5th and 95th (R type-7)
    Pre-registered; trims tails within each meta-analysis.
  • SESOI threshold c = 0.20 (main), 0.10 (sensitivity)
    Pre-registered smallest effect size of interest based on Cohen's 'small' effect.
  • Significance level alpha = 0.05
    Conventional two-sided threshold for statistical significance, pre-registered.
axioms (3)
  • domain assumption The BEAR v2 collection and the k>=10 subsample are representative of behavioral science meta-analyses.
    The analysis covers 259 psychology, 20 psychotherapy, and 79 exercise meta-analyses; the authors explicitly limit scope to behavioral science in Section 5.2, but the generalizability to the wider universe of behavioral science syntheses is an untested assumption.
  • domain assumption The 'do-nothing' baseline represents meta-analyses without prior outlier handling.
    If the source collections already removed or corrected extreme estimates, the measured treatment effects are understated. The authors flag this in Section 5.2 as a limitation.
  • domain assumption REML with Hartung-Knapp-Sidik-Jonkman adjustment and UWLS are appropriate estimators for these pooled effects.
    The paper uses these estimators as standard tools without deriving them; their validity in this context is assumed from the cited literature (e.g., Hartung & Knapp, Stanley et al.).

pith-pipeline@v1.3.0-alltime-deepseek · 14417 in / 11883 out tokens · 106348 ms · 2026-08-01T03:22:48.224623+00:00 · methodology

0 comments
read the original abstract

Meta-analysts routinely face estimates that look too large or extreme. Yet, how to handle them is left to the reviewer's judgment. The methods for detecting such estimates are well known. What is missing is an informed assessment of how much alternative handling choices might change a meta-analysis' conclusions. We fill this gap by analyzing the effects of four pre-registered handling treatments across 358 behavioral science meta-analyses with at least ten estimates. Each outlier handling treatment is estimated by two estimators (random effects and unrestricted weighted least squares), and compared to the 'do-nothing' baseline on three outcomes: the pooled effect, statistical significance, and whether the effect reaches the smallest effect size of interest (|d| >= 0.20). Our entire analysis and comparison pipelines were pre-registered. Alternative outlier handling treatments have little effect on the meta-analysis mean as the median absolute change in Cohen's d is at most 0.047 and often much less. Yet, at least one of these four treatments in combination with one of these estimators reverses the statistical significance of 11.5% of meta-analyses and the smallest-effect-of-interest assessment in 15.9%. Winsorizing has the least effect and DFBETAS the most. Categorical changes are found almost entirely among results already close to the decision boundary; strongly significant results essentially never change. These findings give applied meta-analysts, methods specialists, and reviewers a reference point for how much this under-reported choice matters and provide yet another reason for meta-analysts to publicly pre-specify their methods and handling treatments.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 17 canonical work pages

  1. [1]

    False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant

    Simmons JP, Nelson LD, Simonsohn U. False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science. 2011;22(11):1359-66. doi:10.1177/0956797611417632

  2. [2]

    The Statistical Crisis in Science

    Gelman A, Loken E. The Statistical Crisis in Science. American Scientist. 2014;102(6):460-5. doi:10.1511/2014.111.460

  3. [3]

    Meta-Analyzing the Multiverse: A Peek Under the Hood of Selective Reporting

    Olsson-Collentine A, van Aert RCM, Bakker M, Wicherts JM. Meta-Analyzing the Multiverse: A Peek Under the Hood of Selective Reporting. Psychological Methods. 2025;30(3):441-61. doi:10.1037/met0000559

  4. [4]

    Increasing Transparency Through a Multiverse Analysis

    Steegen S, Tuerlinckx F, Gelman A, Vanpaemel W. Increasing Transparency Through a Multiverse Analysis. Perspectives on Psychological Science. 2016;11(5):702-12. doi:10.1177/1745691616658637

  5. [5]

    Specification Curve Analysis

    Simonsohn U, Simmons JP, Nelson LD. Specification Curve Analysis. Nature Human Behaviour. 2020;4(11):1208-14. doi:10.1038/s41562-020-0912-z

  6. [6]

    Assessment of Vibration of Effects Due to Model Specification Can Demonstrate the Instability of Observational Associations

    Patel CJ, Burford B, Ioannidis JPA. Assessment of Vibration of Effects Due to Model Specification Can Demonstrate the Instability of Observational Associations. Journal of Clinical Epidemiology. 2015;68(9):1046-58. doi:10.1016/j.jclinepi.2015.05.029

  7. [7]

    Many An- alysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results

    Silberzahn R, Uhlmann EL, Martin DP, Anselmi P, Aust F, Awtrey E, et al. Many An- alysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results. Advances in Methods and Practices in Psychological Science. 2018;1(3):337-56. doi:10.1177/2515245917747646

  8. [8]

    A Traveler’s Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions

    Del Giudice M, Gangestad SW. A Traveler’s Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions. Advances in Methods and Practices in Psychological Science. 2021;4(1):1-15. doi:10.1177/2515245920954925

  9. [9]

    Outlier and Influence Diagnostics for Meta-Analysis

    Viechtbauer W, Cheung MW-L. Outlier and Influence Diagnostics for Meta-Analysis. Research Synthesis Methods. 2010;1(2):112-25. doi:10.1002/jrsm.11

  10. [10]

    Sensitivity Analysis in Meta-Analysis: A Tutorial

    Aung NM, Jurak I, Mehmood S, Axon E. Sensitivity Analysis in Meta-Analysis: A Tutorial. Cochrane Evidence Synthesis and Methods. 2026;4(1):e70067. doi:10.1002/cesm.70067

  11. [11]

    Sensitivity Analysis with Iterative Outlier Detection for Systematic Reviews and Meta-Analyses

    Meng Z, Wang J, Lin L, Wu C. Sensitivity Analysis with Iterative Outlier Detection for Systematic Reviews and Meta-Analyses. Statistics in Medicine. 2024;43(8):1549-63. doi:10.1002/sim.10008

  12. [12]

    An Empirical Comparison of Heterogeneity Vari- ance Estimators in 12894 Meta-Analyses

    Langan D, Higgins JPT, Simmonds M. An Empirical Comparison of Heterogeneity Vari- ance Estimators in 12894 Meta-Analyses. Research Synthesis Methods. 2015;6(2):195-205. doi:10.1002/jrsm.1140. 14

  13. [13]

    A Re-Analysis of the Cochrane Library Data: The Dangers of Unobserved Heterogeneity in Meta-Analyses

    Kontopantelis E, Springate DA, Reeves D. A Re-Analysis of the Cochrane Library Data: The Dangers of Unobserved Heterogeneity in Meta-Analyses. PLOS ONE. 2013;8(7):e69930. doi:10.1371/journal.pone.0069930

  14. [14]

    A Novel Robust Meta-Analysis Model Using the t Distribution for Outlier Accommodation and Detection

    Wang Y, Zhao J, Jiang F, Shi L, Pan J. A Novel Robust Meta-Analysis Model Using the t Distribution for Outlier Accommodation and Detection. Research Synthesis Methods. 2025;16(3):442-59. doi:10.1017/rsm.2025.8

  15. [15]

    Robust Inference Methods for Meta-Analysis Involving Influential Outlying Studies

    Noma H, Sugasawa S, Furukawa TA. Robust Inference Methods for Meta-Analysis Involving Influential Outlying Studies. Statistics in Medicine. 2024;43(20):3778-91. doi:10.1002/sim.10157

  16. [16]

    Outlier and Influence Handling in Meta-Analysis: Pre-Analysis Plan; 2026

    Havranek T, Irsova Z, Luskova M, Stanley TD. Outlier and Influence Handling in Meta-Analysis: Pre-Analysis Plan; 2026. https://doi.org/10.17605/OSF.IO/97CMV. Registration, Open Science Framework

  17. [17]

    A Statistical Case for Qualified Scientific Optimism

    van Zwet E, Gelman A, Wiecek W. A Statistical Case for Qualified Scientific Optimism

  18. [18]

    Estimating the Change in Meta-Analytic Effect Size Estimates After the Application of Publication Bias Adjustment Methods

    Sladekova M, Webb LEA, Field AP. Estimating the Change in Meta-Analytic Effect Size Estimates After the Application of Publication Bias Adjustment Methods. Psychological Methods. 2023;28(3):664-86. doi:10.1037/met0000470

  19. [19]

    Exploring the Efficacy of Psychotherapies for Depression: A Multiverse Meta-Analysis

    Plessen CY, Karyotaki E, Miguel C, Ciharova M, Cuijpers P. Exploring the Efficacy of Psychotherapies for Depression: A Multiverse Meta-Analysis. BMJ Mental Health. 2023;26(1):e300626. doi:10.1136/bmjment-2022-300626

  20. [20]

    metapsyData: Access the Meta- Analytic Psychotherapy Databases in R; 2022

    Harrer M, Sprenger AA, Kuper P, Karyotaki E, Cuijpers P. metapsyData: Access the Meta- Analytic Psychotherapy Databases in R; 2022. https://data.metapsy.org. R package, Metapsy Collaboration, Vrije Universiteit Amsterdam

  21. [21]

    Effect of Exercise on Cognition, Memory, and Executive Function: A Study-Level Meta-Meta-Analysis Across Populations and Exercise Categories; 2025

    Bartoš F, Lušková M, Bortnikova K, Hozová K, Kantova K, Irsova Z, Havranek T. Effect of Exercise on Cognition, Memory, and Executive Function: A Study-Level Meta-Meta-Analysis Across Populations and Exercise Categories; 2025. https://doi.org/10.31234/osf.io/ qr8e2_v1. Working paper

  22. [22]

    Conducting Meta-Analyses in R with the metafor Package

    Viechtbauer W. Conducting Meta-Analyses in R with the metafor Package. Journal of Statistical Software. 2010;36(3):1-48. doi:10.18637/jss.v036.i03

  23. [23]

    On Tests of the Overall Treatment Effect in Meta-Analysis with Normally Distributed Responses

    Hartung J, Knapp G. On Tests of the Overall Treatment Effect in Meta-Analysis with Normally Distributed Responses. Statistics in Medicine. 2001;20(12):1771-82. doi:10.1002/sim.791

  24. [24]

    A Simple Confidence Interval for Meta-Analysis

    Sidik K, Jonkman JN. A Simple Confidence Interval for Meta-Analysis. Statistics in Medicine. 2002;21(21):3153-9. doi:10.1002/sim.1262

  25. [25]

    The Hartung-Knapp-Sidik-Jonkman Method for Ran- dom Effects Meta-Analysis Is Straightforward and Considerably Outperforms the Standard DerSimonian-LairdMethod

    IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman Method for Ran- dom Effects Meta-Analysis Is Straightforward and Considerably Outperforms the Standard DerSimonian-LairdMethod. BMCMedicalResearchMethodology.2014;14:25. doi:10.1186/1471- 2288-14-25

  26. [26]

    Neither Fixed nor Random: Weighted Least Squares Meta- Analysis

    Stanley TD, Doucouliagos H. Neither Fixed nor Random: Weighted Least Squares Meta- Analysis. Statistics in Medicine. 2015;34(13):2116-27. doi:10.1002/sim.6481. 15

  27. [27]

    Unre- stricted Weighted Least Squares Represent Medical Research Better than Random Ef- fects in 67,308 Cochrane Meta-Analyses

    Stanley TD, Ioannidis JPA, Maier M, Doucouliagos H, Otte WM, Bartoš F. Unre- stricted Weighted Least Squares Represent Medical Research Better than Random Ef- fects in 67,308 Cochrane Meta-Analyses. Journal of Clinical Epidemiology. 2023;157:53-8. doi:10.1016/j.jclinepi.2023.03.004

  28. [28]

    Why the Unre- stricted Weighted Least Squares Should Be Routinely Reported in Medical Meta-Analyses; 2026.https://osf.io/vpuqj/files/4rkqg

    Stanley TD, Ioannidis JPA, Maier M, Doucouliagos H, Otte WM, Bartoš F. Why the Unre- stricted Weighted Least Squares Should Be Routinely Reported in Medical Meta-Analyses; 2026.https://osf.io/vpuqj/files/4rkqg. Preprint

  29. [29]

    Reducing the Biases of the Conventional Meta- Analysis of Correlations

    Stanley TD, Doucouliagos H, Havranek T. Reducing the Biases of the Conventional Meta- Analysis of Correlations. Research Synthesis Methods. 2025;16(1):42-59. doi:10.1017/rsm.2024.5

  30. [30]

    Statistical Power Analysis for the Behavioral Sciences

    Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988

  31. [31]

    Regression Diagnostics: Identifying Influential Data and Sources of Collinearity

    Belsley DA, Kuh E, Welsch RE. Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. New York: Wiley; 1980. doi:10.1002/0471725153

  32. [32]

    Single-Dataset Meta-Analysis for Many-Analysts and Multiverse Studies; 2025

    Bartoš F, Hoogeveen S, Sarafoglou A, Pawel S. Single-Dataset Meta-Analysis for Many-Analysts and Multiverse Studies; 2025. arXiv:2511.17064 [stat.ME]. https://arxiv.org/abs/2511. 17064. Preprint. 16

  33. [2026]

    Data: Benchmarks of Em- pirical Accuracy in Research (BEAR),https://github.com/wwiecek/BEAR

    Working paper,https://sites.stat.columbia.edu/gelman/research/unpublished/ A_statistical_case_for_qualified_scientific_optimism.pdf. Data: Benchmarks of Em- pirical Accuracy in Research (BEAR),https://github.com/wwiecek/BEAR