REVIEW 3 major objections 5 minor 3 references
In post-publication research ratings, differences between judges explain far more variance than differences between papers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 15:47 UTC pith:2ZJ4LXR6
load-bearing objection Solid large-scale variance partition showing judges dominate papers in H1 Connect ratings; the 61%/7% headline is inflated by same-judge tags, but the core result survives without them. the 3 major comments →
Judges matter more than papers in post-publication research assessment
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In a large post-publication peer-review dataset of quality ratings, judge-related variation (overall severity plus judge-specific weighting of scientific attributes) accounts for substantially more of the variance in ratings than variation attributable to the papers and journals being rated; directional biases such as gender and global affiliation explain almost none of it.
What carries the argument
Multilevel variance partitioning that decomposes ratings into paper intercepts, journal intercepts, judge intercepts (level noise), and judge-specific random slopes on latent factors derived from classification tags (pattern noise).
Load-bearing premise
The classification tags that judges themselves assign can be treated as measures of paper attributes when estimating how differently judges weight those attributes.
What would settle it
A replication in which independent, judge-blind characterizations of the same papers (or an evaluation system with pre-specified fixed criteria) reverse the variance partition so that paper-level effects exceed judge-level effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyses 239,521 H1 Connect post-publication ratings by 12,649 judges of 193,128 papers and partitions rating variance with multilevel models. Baseline random-intercept models attribute more variance to judges (severity) than to papers; adding journals absorbs paper variance but not judge variance. A fuller model adds six latent factors from judge-assigned classification tags as fixed effects and judge-specific random slopes, yielding the headline partition: judge intercepts plus slopes ≈61% of total variance versus paper+journal ≈7%, with demographic/affiliation bias covariates explaining <1%. The authors interpret this as level and pattern noise dominating paper quality differences and argue for routine noise audits in high-stakes research assessment.
Significance. If the core claim holds—that evaluator-related variation substantially exceeds between-paper variation in a real large-scale assessment system—the paper supplies a concrete empirical basis for treating noise audits as standard practice alongside bias checks. Strengths include the unusually large N, progressive model building, and a battery of robustness checks (no-singletons subset, multi-tag subset, permutation collapse of slope variance from ~49% to 0.6%, Bayesian ordinal replication of the intercept partition, multiple optimisers). The bias analyses are a useful negative result. Even a carefully qualified version of the result would be of clear interest to research evaluation, scientometrics, and science policy.
major comments (3)
- [Results: Unmasked pattern noise; Method: Multilevel models] Results “Unmasked pattern noise” and Method (latent-factor construction; formula with (0 + Latent_Factor1+…||judge)): the 49.5% judge-slope term that drives the headline 61%/7% partition is estimated from six latent factors derived from the twelve binary classification tags assigned by the same judges in the same evaluation that produces the rating. Those factors are therefore co-produced with the outcome, not independent measures of paper attributes. The random slopes largely capture within-judge consistency between a judge’s own tags and own rating. The permutation test shows the association is real rather than pure model flexibility, but does not establish that the factors are exogenous paper features. The abstract, Results, and Discussion should not present the 61% figure or the “more than seven times” language as pure pattern noise in the Kahneman sense without a clear qualification
- [Abstract; Results: Judge-level variation exceeded paper-level variation; Discussion] Baseline models without slopes already show judge intercepts exceeding paper intercepts (26% vs 18%; 25% vs ~15% with journals). That qualitative ordering is the load-bearing empirical result and is robust across the no-singletons and Bayesian ordinal checks. The manuscript should lead with, and rest the central claim on, that intercept partition, and treat the slope model as an exploratory decomposition of residual judge structure whose interpretation is limited by co-production. Reframing the abstract and Discussion around the intercept result would make the central claim defensible without overstating the 61%/7% ratio.
- [Discussion] Discussion acknowledges that H1 Connect is a signed recommendation system with expert self-selection of papers, not gatekeeping peer review, and notes the compressed quality range. That selection and range restriction is still under-analysed as a threat to the paper-level variance component: if experts only rate papers they already view as above a high threshold, between-paper variance is mechanically reduced and the judge/paper ratio is inflated. A quantitative sensitivity discussion (or bounds) on how much paper variance could be missing under plausible selection would strengthen the claim that judges dominate papers rather than that the platform samples a narrow quality band.
minor comments (5)
- [Figure 1] Figure 1 caption and text: report exact variance percentages for each bar segment (or a companion table) so readers can verify the 61%/7% and robustness comparisons without estimating from the stacked bars.
- [Method: Robustness analyses] Method: “no singletons” case counts are given as both 76,698 and 76,398 in adjacent paragraphs; reconcile the figure.
- [Abstract; Results; Discussion] Typos: “paritioning” (Discussion), “difference in the evaluated research” (Abstract, should be “differences”), and occasional “rankings” where “ratings” is meant (Results, journal model).
- [Method: The H1 Connect dataset] Clarify whether latent factor scores are paper-level aggregates or evaluation-level (judge×paper) scores; the text says “for each evaluated paper” but tags are assigned per evaluation. This affects interpretation of the random slopes.
- [Appendix] Appendix bias models: state sample sizes after gender_guesser exclusions more prominently so the <1% variance claim is easy to locate next to the main partition.
Circularity Check
Empirical variance partition is self-contained; only mild interpretive circularity arises from treating same-judge tags as exogenous paper attributes for the 49.5% slope term.
specific steps
-
other
[Results: Unmasked pattern noise; Method: latent-factor construction and model formula with (0 + Latent_Factor1+…||judge)]
"We incorporated these latent factors into the previous model by including them as fixed effects and by fitting judge-specific random slopes for each factor. … Variance partitioning showed that judge-level effects (11.5%) plus the sum of judge-slope effects (49.5%) accounted for 61% of total variance in ratings, as compared to 7% for combined paper and journal-level effects. … Latent factor scores were calculated for each evaluated paper and used as predictors in subsequent multilevel models."
The twelve binary classification tags (and the six latent factors derived from them) are assigned by the same judges in the same evaluation act that produces the rating. The random slopes therefore largely recover within-judge consistency between a judge’s own tags and own rating, not differential weighting of exogenous paper attributes. Calling the factors “latent paper dimensions” or “paper characteristics” imports the co-produced structure into the “pattern-noise” term by construction of the predictors, inflating the 61% figure that drives the strongest claim.
full rationale
The paper reports a descriptive multilevel variance decomposition of observed H1 Connect ratings; the numerical partitions (including the baseline 26% judge vs 18% paper intercepts) are obtained directly from the fitted random-effects models and do not reduce by construction to any quantity defined in terms of a fitted target or self-cited uniqueness result. Latent factors are data-driven via parallel analysis, and the permutation test correctly shows that slope variance is not an artefact of model flexibility. Self-citations (Ward 2026; Bornmann reviews) supply background and implications only. The sole soft link is interpretive: the six latent factors used for judge-specific slopes are derived from the twelve binary classification tags that the same judges assign in the same evaluation that produces the rating. Consequently the large (49.5%) slope contribution partly captures within-judge tag–rating consistency rather than differential weighting of independently measured paper attributes. This does not make the reported percentages circular by equation, but it does inflate the headline 61%/7% contrast relative to a pure paper-attribute interpretation. Baseline models without slopes already show judge > paper variance, so the central claim retains independent empirical content. Score therefore remains low (2).
Axiom & Free-Parameter Ledger
free parameters (1)
- number of latent factors from classification tags =
6
axioms (3)
- standard math Linear mixed-effects models with REML (and a Bayesian cumulative ordinal counterpart) correctly partition rating variance into paper, journal, judge-intercept, and judge-slope components.
- domain assumption H1 Connect signed recommendations of already-published biomedical papers are informative about the relative size of judge versus paper variance in high-stakes research assessment more generally.
- ad hoc to paper Judge-assigned classification tags can be treated as measures of paper attributes for estimating pattern noise.
read the original abstract
Research assessment relies on expert evaluations, yet human judgement is noisy, and it is unclear whether differences in assessment arise primarily from differences in genuine research quality or from unwanted differences between evaluators. While numerous studies highlight disagreement and biases in research assessment, they have not quantified judge-related noise relative to variation in the evaluated works. Here we show, in a large post-publication peer review database, that research assessment is driven more by differences between evaluators than by difference in the evaluated research. We partition variance in 239,521 research quality ratings assigned by 12,649 judges to 193,128 papers from the H1 Connect post-publication peer review platform. Using multilevel models, we decomposed judge-related variation into differences in overall severity and differences in the weighting of scientific attributes. We found that judge-related effects accounted for substantially more variance in ratings than the evaluated papers. In our most detailed model, judge-level effects and judge-specific slopes explained 61% of the total variance, whereas combined paper and journal-level effects accounted for only 7%. By contrast, examined measures of directional bias, such as author gender and global affiliation, explained less than 1% of the variance. We conclude that assessment outcomes were shaped more by the judges than by the papers themselves. Our results demonstrate the necessity of noise audits in high-stakes scientific evaluation.
Figures
Reference graph
Works this paper leans on
-
[1]
H., Reisch, L
Allison, K. H., Reisch, L. M., Carney, P. A., Weaver, D. L., Schnitt, S. J., O’Malley, F. P., & Elmore, J. G. (2014). Understanding diagnostic variability in breast pathology: Lessons learned from an expert consensus review panel. Histopathology, 65(2), 240–
2014
-
[2]
Bonavia, T., & Marin-Garcia, J. A. (2023). A noise audit of the peer review of a scientific article: A WPOM journal case study. WPOM-Working Papers on Operations Management, 14(2), 137-166. https://doi.org/10.4995/wpom.19631 Bornmann, L. (2011). Scientific peer review. Annual Review of Information Science and Technology, 45, 199-245. https://doi.org/10.10...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.4995/wpom.19631 2023
-
[3]
Murray, D., Siler, K., Larivière, V., Chan, W. M., Collings, A. M., Raymond, J., & Sugimoto, C. R. (2019). Author-reviewer homophily in peer review. Retrieved January 23, 2026, from https://www.biorxiv.org/content/biorxiv/early/2019/08/04/400515.full.pdf Pier, E. L., Brauer, M., Filut, A., Kaatz, A., Raclaw, J., Nathan, M. J., & Carnes, M. (2018). Low agr...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.