REVIEW 3 minor 33 references
Outlier-handling choices flip the verdict in 11.5% of behavioral science meta-analyses.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:22 UTC pith:NOR6YTPH
load-bearing objection Solid, pre-registered empirical benchmark: first large-scale comparison of outlier-handling treatments in meta-analysis; the one soft spot (baseline integrity) is disclosed and likely conservative.
Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that in behavioral science meta-analyses with at least ten estimates, outlier and influence handling decisions—dropping the most extreme estimate, deleting estimates flagged by studentized deleted residuals or DFBETAS, or winsorizing the tails—have little effect on the magnitude of the pooled mean (median |Δd| ≤ 0.047), but a meaningful effect on categorical interpretation. Counting either of two estimators and any of the four treatments, 7.7% of the 715 estimable estimator-by-meta-analysis cells change statistical significance and 10.3% change smallest-effect-of-interest status; at the level of meta-analyses the rates are 11.5% and 15.9%. The authors emphasize that wins
What carries the argument
The comparison machinery is a pre-registered grid: each of 358 meta-analyses is estimated under two estimators—random-effects with a small-sample t correction, and unrestricted weighted least squares (an inverse-variance weighted mean with heterogeneity-inflated standard errors)—and each is compared against a 'do nothing' baseline on four treatments: dropping the single most extreme effect, deleting estimates with studentized deleted residuals above 3, winsorizing at the 5th and 95th percentiles, and deleting estimates with |DFBETAS| above 2/√k. Treating one handling choice at a time while holding thresholds and edge rules fixed makes any observed change attributable to the handling rule. Th
Load-bearing premise
The 'do nothing' baseline is assumed to reflect genuinely unhandled data; if the source meta-analyses had already removed or corrected extreme estimates, the measured reversal rates understate how much outlier handling matters.
What would settle it
One finding that would settle it: re-run the pre-registered treatments on a set of meta-analyses where raw primary-study data are available and the original authors' outlier removals are documented, and build the do-nothing baseline from the unhandled raw estimates. If reversal rates fall to near zero, the reported 11.5% and 15.9% figures are an artifact of baselines that already incorporate outlier handling. Conversely, if reversals were found concentrated among strongly significant or large effects—which the paper claims essentially never change—the central claim would also be contradicted.
If this is right
- Meta-analysts should pre-specify their outlier handling rule: a defensible choice of method can change the headline verdict in about one in nine syntheses.
- Reviewers should ask for sensitivity analyses that report the central result both with and without outlier treatment; borderline findings are where the risk of reversal lives.
- Winsorizing is the least disruptive handling option and DFBETAS-based deletion the most; a meta-analyst wanting a conservative check can use winsorizing, while DFBETAS under unrestricted weighted least squares is the most sensitive diagnostic to report.
- Because reversal rates rise when meta-analyses are small (k below 20) and fall when they are large, outlier handling deserves special attention in small syntheses.
- These rates are a benchmark: a 7.7% cell-level and 11.5% meta-analysis-level significance-reversal rate is comparable in scale to the discordance previously documented for heterogeneity-estimator choice, so outlier handling is no smaller a degree of freedom.
Where Pith is reading between the lines
- The reversal rates are measured against a 'do nothing' baseline taken from compiled meta-analyses that may already have had outliers removed by their original authors; if so, the true sensitivity of meta-analytic conclusions to outlier handling is understated, not overstated.
- Because the sample is limited to behavioral science syntheses with at least ten effects, applying the same pre-registered pipeline to medical or other disciplinary meta-analyses, where effect distributions and typical k differ, could meaningfully change the rates; that is a testable extension of the design.
- The fact that winsorizing only ever moves results toward statistical significance suggests a caution: a method that shrinks tails without asking why an estimate is extreme could make a synthesis look more conclusive than the data warrant, especially in areas with publication bias.
- A direct practical test of this design's value: for a set of meta-analyses with fully documented raw data, compare the conclusions reached by teams that pre-registered versus teams that did not, to see whether pre-specification actually reduces outcome-dependent outlier handling.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper quantifies how pre-registered outlier and influence-handling decisions affect meta-analytic conclusions in 358 behavioral-science meta-analyses (k≥10; 16,664 estimates) from the BEAR v2 collection spanning psychology, psychotherapy, and exercise. Four active treatments—dropping the most extreme |d|, deleting estimates with studentized deleted residuals >3, winsorizing at the 5th/95th percentiles, and deleting estimates with |DFBETAS|>2/√k—are compared against a do-nothing baseline under two estimators: random effects (REML with Hartung–Knapp adjustment) and unrestricted weighted least squares (UWLS). Outcomes are absolute change in pooled d, statistical significance (p<.05), and whether |pooled d| reaches a SESOI of ≥0.20. The median absolute change in d is at most 0.047 across treatments. At least one treatment changes statistical significance in 7.7% of 715 estimable estimator-by-meta-analysis cells (11.5% of 358 meta-analyses) and changes SESOI status in 10.3% of cells (15.9% of meta-analyses). Reversals are largely confined to results near the decision boundary; DFBETAS is the most sensitive and winsorizing the least. Robustness checks (k≥20, k≥5, alternative cutoffs, one-estimate-per-study) are consistent. The entire pipeline was pre-registered, the dataset is frozen, and a replication package is provided.
Significance. If correct, this is the first large-scale empirical reference for the sensitivity of meta-analytic significance and substantive-size verdicts to outlier handling, an under-reported researcher degree of freedom. The paper's strengths are its design: a pre-analysis plan with fixed thresholds and edge rules, a frozen dataset with checksum, and a publicly archived replication package that makes the confirmatory results push-button reproducible. The comparison to Langan et al.'s heterogeneity-estimator discordance (7.5%) provides a useful benchmark. The descriptive nature is appropriate; the authors do not claim to identify which treatment is correct. Limitations—possible prior cleaning in the source databases, repeated d/SE rows of unknown provenance in psychology, and the conflation of small-study effects with true outliers—are acknowledged and partially mitigated by the retention of extreme values (|d|>10) in the psychology source. The findings should be of interest to applied meta-analysts, methods researchers, and journal reviewers.
minor comments (3)
- [§5.2, Limitations] The statement that prior cleaning in the source databases would make the treatment effects "understate how much these choices matter" is plausible but not logically guaranteed: if earlier cleaning selectively removed estimates that were propping up significance, the remaining baseline could be more fragile, and the reported reversal rates could overstate sensitivity. The retention of 37 estimates with |d|>10 supports the conservative-direction reading, but a sentence acknowledging the alternative direction and noting that a direct test would require access to source pre-processing would be more precise.
- [Table 4, notes] The column "Not estimable" in Table 4 is not defined in the table notes. The text in §4.4 explains that these arise from metafor REML non-convergence, but the table note should carry this definition for readers who start with the tables.
- [§3.3, Outcomes] The SESOI threshold c=0.20 is based on Cohen's small effect for d. Because the psychology studies are converted from Fisher-z correlations, it may be worth adding a sentence clarifying whether the threshold is intended to be directly comparable across the three fields, which have different original metrics.
Circularity Check
No circularity: the paper reports preregistered measurements on a fixed dataset, not predictions derived from a fitted model.
full rationale
The paper's central results are direct empirical summaries: it applies four fixed outlier-handling treatments and two estimators to a frozen dataset, counts how often the statistical-significance or SESOI status changes relative to a do-nothing baseline, and reports median absolute changes in Cohen's d. There is no estimation-then-prediction loop: every treatment, threshold, estimator, and outcome was fixed in the pre-analysis plan before outcomes were computed, so the reversal rates cannot be a product of post hoc fitting. The estimators (RE with HKSJ and UWLS) are defined explicitly in Section 3.2, and UWLS is justified by the Gauss-Markov theorem rather than by an imported uniqueness claim. Although some citations to the authors' own prior work motivate UWLS and supply one of the three source collections, these are not load-bearing in the sense of forcing the result; the exercise-data overlap is disclosed in the Competing Interest statement and affects sample composition, not the derivation of the measured statistics. The Discussion's statement that DFBETAS should have the most effect 'because this is what DFBETAS measures' is an acknowledged expectation from the definition of the diagnostic, not a predicted result that is then treated as independently derived. The Section 5.2 limitation that the 'do-nothing' baseline may already contain prior outlier-handling decisions is a measurement/validity concern about the interpretation of the reversal rates; it does not make any reported rate equivalent to an input by construction. Overall, the paper is a self-contained empirical sensitivity survey with preregistered comparisons, and no circular step is exhibited.
Axiom & Free-Parameter Ledger
free parameters (5)
- Studentized deleted residual cutoff =
3
- DFBETAS deletion threshold =
2/sqrt(k)
- Winsorizing percentiles =
5th and 95th (R type-7)
- SESOI threshold c =
0.20 (main), 0.10 (sensitivity)
- Significance level alpha =
0.05
axioms (3)
- domain assumption The BEAR v2 collection and the k>=10 subsample are representative of behavioral science meta-analyses.
- domain assumption The 'do-nothing' baseline represents meta-analyses without prior outlier handling.
- domain assumption REML with Hartung-Knapp-Sidik-Jonkman adjustment and UWLS are appropriate estimators for these pooled effects.
read the original abstract
Meta-analysts routinely face estimates that look too large or extreme. Yet, how to handle them is left to the reviewer's judgment. The methods for detecting such estimates are well known. What is missing is an informed assessment of how much alternative handling choices might change a meta-analysis' conclusions. We fill this gap by analyzing the effects of four pre-registered handling treatments across 358 behavioral science meta-analyses with at least ten estimates. Each outlier handling treatment is estimated by two estimators (random effects and unrestricted weighted least squares), and compared to the 'do-nothing' baseline on three outcomes: the pooled effect, statistical significance, and whether the effect reaches the smallest effect size of interest (|d| >= 0.20). Our entire analysis and comparison pipelines were pre-registered. Alternative outlier handling treatments have little effect on the meta-analysis mean as the median absolute change in Cohen's d is at most 0.047 and often much less. Yet, at least one of these four treatments in combination with one of these estimators reverses the statistical significance of 11.5% of meta-analyses and the smallest-effect-of-interest assessment in 15.9%. Winsorizing has the least effect and DFBETAS the most. Categorical changes are found almost entirely among results already close to the decision boundary; strongly significant results essentially never change. These findings give applied meta-analysts, methods specialists, and reviewers a reference point for how much this under-reported choice matters and provide yet another reason for meta-analysts to publicly pre-specify their methods and handling treatments.
Reference graph
Works this paper leans on
-
[1]
Simmons JP, Nelson LD, Simonsohn U. False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science. 2011;22(11):1359-66. doi:10.1177/0956797611417632
-
[2]
The Statistical Crisis in Science
Gelman A, Loken E. The Statistical Crisis in Science. American Scientist. 2014;102(6):460-5. doi:10.1511/2014.111.460
-
[3]
Meta-Analyzing the Multiverse: A Peek Under the Hood of Selective Reporting
Olsson-Collentine A, van Aert RCM, Bakker M, Wicherts JM. Meta-Analyzing the Multiverse: A Peek Under the Hood of Selective Reporting. Psychological Methods. 2025;30(3):441-61. doi:10.1037/met0000559
-
[4]
Increasing Transparency Through a Multiverse Analysis
Steegen S, Tuerlinckx F, Gelman A, Vanpaemel W. Increasing Transparency Through a Multiverse Analysis. Perspectives on Psychological Science. 2016;11(5):702-12. doi:10.1177/1745691616658637
-
[5]
Simonsohn U, Simmons JP, Nelson LD. Specification Curve Analysis. Nature Human Behaviour. 2020;4(11):1208-14. doi:10.1038/s41562-020-0912-z
-
[6]
Patel CJ, Burford B, Ioannidis JPA. Assessment of Vibration of Effects Due to Model Specification Can Demonstrate the Instability of Observational Associations. Journal of Clinical Epidemiology. 2015;68(9):1046-58. doi:10.1016/j.jclinepi.2015.05.029
-
[7]
Many An- alysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results
Silberzahn R, Uhlmann EL, Martin DP, Anselmi P, Aust F, Awtrey E, et al. Many An- alysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results. Advances in Methods and Practices in Psychological Science. 2018;1(3):337-56. doi:10.1177/2515245917747646
-
[8]
Del Giudice M, Gangestad SW. A Traveler’s Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions. Advances in Methods and Practices in Psychological Science. 2021;4(1):1-15. doi:10.1177/2515245920954925
-
[9]
Outlier and Influence Diagnostics for Meta-Analysis
Viechtbauer W, Cheung MW-L. Outlier and Influence Diagnostics for Meta-Analysis. Research Synthesis Methods. 2010;1(2):112-25. doi:10.1002/jrsm.11
doi:10.1002/jrsm.11 2010
-
[10]
Sensitivity Analysis in Meta-Analysis: A Tutorial
Aung NM, Jurak I, Mehmood S, Axon E. Sensitivity Analysis in Meta-Analysis: A Tutorial. Cochrane Evidence Synthesis and Methods. 2026;4(1):e70067. doi:10.1002/cesm.70067
-
[11]
Sensitivity Analysis with Iterative Outlier Detection for Systematic Reviews and Meta-Analyses
Meng Z, Wang J, Lin L, Wu C. Sensitivity Analysis with Iterative Outlier Detection for Systematic Reviews and Meta-Analyses. Statistics in Medicine. 2024;43(8):1549-63. doi:10.1002/sim.10008
-
[12]
An Empirical Comparison of Heterogeneity Vari- ance Estimators in 12894 Meta-Analyses
Langan D, Higgins JPT, Simmonds M. An Empirical Comparison of Heterogeneity Vari- ance Estimators in 12894 Meta-Analyses. Research Synthesis Methods. 2015;6(2):195-205. doi:10.1002/jrsm.1140. 14
-
[13]
A Re-Analysis of the Cochrane Library Data: The Dangers of Unobserved Heterogeneity in Meta-Analyses
Kontopantelis E, Springate DA, Reeves D. A Re-Analysis of the Cochrane Library Data: The Dangers of Unobserved Heterogeneity in Meta-Analyses. PLOS ONE. 2013;8(7):e69930. doi:10.1371/journal.pone.0069930
-
[14]
A Novel Robust Meta-Analysis Model Using the t Distribution for Outlier Accommodation and Detection
Wang Y, Zhao J, Jiang F, Shi L, Pan J. A Novel Robust Meta-Analysis Model Using the t Distribution for Outlier Accommodation and Detection. Research Synthesis Methods. 2025;16(3):442-59. doi:10.1017/rsm.2025.8
-
[15]
Robust Inference Methods for Meta-Analysis Involving Influential Outlying Studies
Noma H, Sugasawa S, Furukawa TA. Robust Inference Methods for Meta-Analysis Involving Influential Outlying Studies. Statistics in Medicine. 2024;43(20):3778-91. doi:10.1002/sim.10157
-
[16]
Outlier and Influence Handling in Meta-Analysis: Pre-Analysis Plan; 2026
Havranek T, Irsova Z, Luskova M, Stanley TD. Outlier and Influence Handling in Meta-Analysis: Pre-Analysis Plan; 2026. https://doi.org/10.17605/OSF.IO/97CMV. Registration, Open Science Framework
-
[17]
A Statistical Case for Qualified Scientific Optimism
van Zwet E, Gelman A, Wiecek W. A Statistical Case for Qualified Scientific Optimism
-
[18]
Sladekova M, Webb LEA, Field AP. Estimating the Change in Meta-Analytic Effect Size Estimates After the Application of Publication Bias Adjustment Methods. Psychological Methods. 2023;28(3):664-86. doi:10.1037/met0000470
-
[19]
Exploring the Efficacy of Psychotherapies for Depression: A Multiverse Meta-Analysis
Plessen CY, Karyotaki E, Miguel C, Ciharova M, Cuijpers P. Exploring the Efficacy of Psychotherapies for Depression: A Multiverse Meta-Analysis. BMJ Mental Health. 2023;26(1):e300626. doi:10.1136/bmjment-2022-300626
-
[20]
metapsyData: Access the Meta- Analytic Psychotherapy Databases in R; 2022
Harrer M, Sprenger AA, Kuper P, Karyotaki E, Cuijpers P. metapsyData: Access the Meta- Analytic Psychotherapy Databases in R; 2022. https://data.metapsy.org. R package, Metapsy Collaboration, Vrije Universiteit Amsterdam
2022
-
[21]
Bartoš F, Lušková M, Bortnikova K, Hozová K, Kantova K, Irsova Z, Havranek T. Effect of Exercise on Cognition, Memory, and Executive Function: A Study-Level Meta-Meta-Analysis Across Populations and Exercise Categories; 2025. https://doi.org/10.31234/osf.io/ qr8e2_v1. Working paper
doi:10.31234/osf.io/ 2025
-
[22]
Conducting Meta-Analyses in R with the metafor Package
Viechtbauer W. Conducting Meta-Analyses in R with the metafor Package. Journal of Statistical Software. 2010;36(3):1-48. doi:10.18637/jss.v036.i03
-
[23]
On Tests of the Overall Treatment Effect in Meta-Analysis with Normally Distributed Responses
Hartung J, Knapp G. On Tests of the Overall Treatment Effect in Meta-Analysis with Normally Distributed Responses. Statistics in Medicine. 2001;20(12):1771-82. doi:10.1002/sim.791
-
[24]
A Simple Confidence Interval for Meta-Analysis
Sidik K, Jonkman JN. A Simple Confidence Interval for Meta-Analysis. Statistics in Medicine. 2002;21(21):3153-9. doi:10.1002/sim.1262
-
[25]
IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman Method for Ran- dom Effects Meta-Analysis Is Straightforward and Considerably Outperforms the Standard DerSimonian-LairdMethod. BMCMedicalResearchMethodology.2014;14:25. doi:10.1186/1471- 2288-14-25
doi:10.1186/1471- 2014
-
[26]
Neither Fixed nor Random: Weighted Least Squares Meta- Analysis
Stanley TD, Doucouliagos H. Neither Fixed nor Random: Weighted Least Squares Meta- Analysis. Statistics in Medicine. 2015;34(13):2116-27. doi:10.1002/sim.6481. 15
-
[27]
Stanley TD, Ioannidis JPA, Maier M, Doucouliagos H, Otte WM, Bartoš F. Unre- stricted Weighted Least Squares Represent Medical Research Better than Random Ef- fects in 67,308 Cochrane Meta-Analyses. Journal of Clinical Epidemiology. 2023;157:53-8. doi:10.1016/j.jclinepi.2023.03.004
-
[28]
Why the Unre- stricted Weighted Least Squares Should Be Routinely Reported in Medical Meta-Analyses; 2026.https://osf.io/vpuqj/files/4rkqg
Stanley TD, Ioannidis JPA, Maier M, Doucouliagos H, Otte WM, Bartoš F. Why the Unre- stricted Weighted Least Squares Should Be Routinely Reported in Medical Meta-Analyses; 2026.https://osf.io/vpuqj/files/4rkqg. Preprint
2026
-
[29]
Reducing the Biases of the Conventional Meta- Analysis of Correlations
Stanley TD, Doucouliagos H, Havranek T. Reducing the Biases of the Conventional Meta- Analysis of Correlations. Research Synthesis Methods. 2025;16(1):42-59. doi:10.1017/rsm.2024.5
-
[30]
Statistical Power Analysis for the Behavioral Sciences
Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988
1988
-
[31]
Regression Diagnostics: Identifying Influential Data and Sources of Collinearity
Belsley DA, Kuh E, Welsch RE. Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. New York: Wiley; 1980. doi:10.1002/0471725153
-
[32]
Single-Dataset Meta-Analysis for Many-Analysts and Multiverse Studies; 2025
Bartoš F, Hoogeveen S, Sarafoglou A, Pawel S. Single-Dataset Meta-Analysis for Many-Analysts and Multiverse Studies; 2025. arXiv:2511.17064 [stat.ME]. https://arxiv.org/abs/2511. 17064. Preprint. 16
arXiv 2025
-
[2026]
Data: Benchmarks of Em- pirical Accuracy in Research (BEAR),https://github.com/wwiecek/BEAR
Working paper,https://sites.stat.columbia.edu/gelman/research/unpublished/ A_statistical_case_for_qualified_scientific_optimism.pdf. Data: Benchmarks of Em- pirical Accuracy in Research (BEAR),https://github.com/wwiecek/BEAR
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.