Pith. sign in

REVIEW 3 major objections 5 minor 3 references

In post-publication research ratings, differences between judges explain far more variance than differences between papers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:47 UTC pith:2ZJ4LXR6

load-bearing objection Solid large-scale variance partition showing judges dominate papers in H1 Connect ratings; the 61%/7% headline is inflated by same-judge tags, but the core result survives without them. the 3 major comments →

arxiv 2607.09783 v1 pith:2ZJ4LXR6 submitted 2026-07-08 cs.DL

Judges matter more than papers in post-publication research assessment

classification cs.DL
keywords Research assessmentPeer reviewEvaluator noiseVariance partitioningMultilevel modellingSystematic biasPost-publication peer review
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Expert research assessment is meant to track the quality of the work under review, yet human judgment is noisy. This paper asks which source of variation is larger: genuine differences among papers, or systematic differences among the people who rate them. Using multilevel models on hundreds of thousands of signed quality ratings from a large biomedical post-publication review platform, the authors partition the variance into paper effects, journal effects, judge severity, and judge-specific ways of weighting scientific attributes. Across models, judge-related components dominate: in the fullest specification they account for roughly 61 percent of total variance, while paper and journal effects together account for only about 7 percent. Demographic and affiliation-related biases explain less than 1 percent. The authors conclude that assessment outcomes are shaped more by who is judging than by what is being judged, and that routine noise audits that separate judge variance from paper variance should become standard in high-stakes evaluation.

Core claim

In a large post-publication peer-review dataset of quality ratings, judge-related variation (overall severity plus judge-specific weighting of scientific attributes) accounts for substantially more of the variance in ratings than variation attributable to the papers and journals being rated; directional biases such as gender and global affiliation explain almost none of it.

What carries the argument

Multilevel variance partitioning that decomposes ratings into paper intercepts, journal intercepts, judge intercepts (level noise), and judge-specific random slopes on latent factors derived from classification tags (pattern noise).

Load-bearing premise

The classification tags that judges themselves assign can be treated as measures of paper attributes when estimating how differently judges weight those attributes.

What would settle it

A replication in which independent, judge-blind characterizations of the same papers (or an evaluation system with pre-specified fixed criteria) reverse the variance partition so that paper-level effects exceed judge-level effects.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper analyses 239,521 H1 Connect post-publication ratings by 12,649 judges of 193,128 papers and partitions rating variance with multilevel models. Baseline random-intercept models attribute more variance to judges (severity) than to papers; adding journals absorbs paper variance but not judge variance. A fuller model adds six latent factors from judge-assigned classification tags as fixed effects and judge-specific random slopes, yielding the headline partition: judge intercepts plus slopes ≈61% of total variance versus paper+journal ≈7%, with demographic/affiliation bias covariates explaining <1%. The authors interpret this as level and pattern noise dominating paper quality differences and argue for routine noise audits in high-stakes research assessment.

Significance. If the core claim holds—that evaluator-related variation substantially exceeds between-paper variation in a real large-scale assessment system—the paper supplies a concrete empirical basis for treating noise audits as standard practice alongside bias checks. Strengths include the unusually large N, progressive model building, and a battery of robustness checks (no-singletons subset, multi-tag subset, permutation collapse of slope variance from ~49% to 0.6%, Bayesian ordinal replication of the intercept partition, multiple optimisers). The bias analyses are a useful negative result. Even a carefully qualified version of the result would be of clear interest to research evaluation, scientometrics, and science policy.

major comments (3)
  1. [Results: Unmasked pattern noise; Method: Multilevel models] Results “Unmasked pattern noise” and Method (latent-factor construction; formula with (0 + Latent_Factor1+…||judge)): the 49.5% judge-slope term that drives the headline 61%/7% partition is estimated from six latent factors derived from the twelve binary classification tags assigned by the same judges in the same evaluation that produces the rating. Those factors are therefore co-produced with the outcome, not independent measures of paper attributes. The random slopes largely capture within-judge consistency between a judge’s own tags and own rating. The permutation test shows the association is real rather than pure model flexibility, but does not establish that the factors are exogenous paper features. The abstract, Results, and Discussion should not present the 61% figure or the “more than seven times” language as pure pattern noise in the Kahneman sense without a clear qualification
  2. [Abstract; Results: Judge-level variation exceeded paper-level variation; Discussion] Baseline models without slopes already show judge intercepts exceeding paper intercepts (26% vs 18%; 25% vs ~15% with journals). That qualitative ordering is the load-bearing empirical result and is robust across the no-singletons and Bayesian ordinal checks. The manuscript should lead with, and rest the central claim on, that intercept partition, and treat the slope model as an exploratory decomposition of residual judge structure whose interpretation is limited by co-production. Reframing the abstract and Discussion around the intercept result would make the central claim defensible without overstating the 61%/7% ratio.
  3. [Discussion] Discussion acknowledges that H1 Connect is a signed recommendation system with expert self-selection of papers, not gatekeeping peer review, and notes the compressed quality range. That selection and range restriction is still under-analysed as a threat to the paper-level variance component: if experts only rate papers they already view as above a high threshold, between-paper variance is mechanically reduced and the judge/paper ratio is inflated. A quantitative sensitivity discussion (or bounds) on how much paper variance could be missing under plausible selection would strengthen the claim that judges dominate papers rather than that the platform samples a narrow quality band.
minor comments (5)
  1. [Figure 1] Figure 1 caption and text: report exact variance percentages for each bar segment (or a companion table) so readers can verify the 61%/7% and robustness comparisons without estimating from the stacked bars.
  2. [Method: Robustness analyses] Method: “no singletons” case counts are given as both 76,698 and 76,398 in adjacent paragraphs; reconcile the figure.
  3. [Abstract; Results; Discussion] Typos: “paritioning” (Discussion), “difference in the evaluated research” (Abstract, should be “differences”), and occasional “rankings” where “ratings” is meant (Results, journal model).
  4. [Method: The H1 Connect dataset] Clarify whether latent factor scores are paper-level aggregates or evaluation-level (judge×paper) scores; the text says “for each evaluated paper” but tags are assigned per evaluation. This affects interpretation of the random slopes.
  5. [Appendix] Appendix bias models: state sample sizes after gender_guesser exclusions more prominently so the <1% variance claim is easy to locate next to the main partition.

Circularity Check

1 steps flagged

Empirical variance partition is self-contained; only mild interpretive circularity arises from treating same-judge tags as exogenous paper attributes for the 49.5% slope term.

specific steps
  1. other [Results: Unmasked pattern noise; Method: latent-factor construction and model formula with (0 + Latent_Factor1+…||judge)]
    "We incorporated these latent factors into the previous model by including them as fixed effects and by fitting judge-specific random slopes for each factor. … Variance partitioning showed that judge-level effects (11.5%) plus the sum of judge-slope effects (49.5%) accounted for 61% of total variance in ratings, as compared to 7% for combined paper and journal-level effects. … Latent factor scores were calculated for each evaluated paper and used as predictors in subsequent multilevel models."

    The twelve binary classification tags (and the six latent factors derived from them) are assigned by the same judges in the same evaluation act that produces the rating. The random slopes therefore largely recover within-judge consistency between a judge’s own tags and own rating, not differential weighting of exogenous paper attributes. Calling the factors “latent paper dimensions” or “paper characteristics” imports the co-produced structure into the “pattern-noise” term by construction of the predictors, inflating the 61% figure that drives the strongest claim.

full rationale

The paper reports a descriptive multilevel variance decomposition of observed H1 Connect ratings; the numerical partitions (including the baseline 26% judge vs 18% paper intercepts) are obtained directly from the fitted random-effects models and do not reduce by construction to any quantity defined in terms of a fitted target or self-cited uniqueness result. Latent factors are data-driven via parallel analysis, and the permutation test correctly shows that slope variance is not an artefact of model flexibility. Self-citations (Ward 2026; Bornmann reviews) supply background and implications only. The sole soft link is interpretive: the six latent factors used for judge-specific slopes are derived from the twelve binary classification tags that the same judges assign in the same evaluation that produces the rating. Consequently the large (49.5%) slope contribution partly captures within-judge tag–rating consistency rather than differential weighting of independently measured paper attributes. This does not make the reported percentages circular by equation, but it does inflate the headline 61%/7% contrast relative to a pure paper-attribute interpretation. Baseline models without slopes already show judge > paper variance, so the central claim retains independent empirical content. Score therefore remains low (2).

Axiom & Free-Parameter Ledger

1 free parameters · 3 axioms · 0 invented entities

The central claim is an observational variance partition. It rests on standard multilevel-model assumptions, the H1 Connect design (signed recommendations of already-published biomedical papers), and a data-driven reduction of twelve binary tags to six latent factors. No new physical entities or free physical constants are introduced; the main free modelling choices are the number of factors and the continuous-versus-ordinal treatment of the three-level rating (both checked).

free parameters (1)
  • number of latent factors from classification tags = 6
    Parallel analysis selected six factors explaining 65% of tag-usage variance; these become the predictors that receive judge-specific random slopes and drive the 49.5% pattern-noise component.
axioms (3)
  • standard math Linear mixed-effects models with REML (and a Bayesian cumulative ordinal counterpart) correctly partition rating variance into paper, journal, judge-intercept, and judge-slope components.
    Invoked throughout the Method and Results; standard in multilevel modelling but assumes the random-effects structure matches the data-generating process.
  • domain assumption H1 Connect signed recommendations of already-published biomedical papers are informative about the relative size of judge versus paper variance in high-stakes research assessment more generally.
    Stated in Discussion when linking results to the UK REF and gatekeeping peer review; the sample is restricted to papers experts chose to recommend.
  • ad hoc to paper Judge-assigned classification tags can be treated as measures of paper attributes for estimating pattern noise.
    Used to construct the six latent factors and the random-slope terms that account for most of the 61% judge-related variance; tags and ratings are produced by the same judge in the same evaluation.

pith-pipeline@v1.1.0-grok45 · 15084 in / 2835 out tokens · 47751 ms · 2026-07-14T15:47:28.237784+00:00 · methodology

0 comments
read the original abstract

Research assessment relies on expert evaluations, yet human judgement is noisy, and it is unclear whether differences in assessment arise primarily from differences in genuine research quality or from unwanted differences between evaluators. While numerous studies highlight disagreement and biases in research assessment, they have not quantified judge-related noise relative to variation in the evaluated works. Here we show, in a large post-publication peer review database, that research assessment is driven more by differences between evaluators than by difference in the evaluated research. We partition variance in 239,521 research quality ratings assigned by 12,649 judges to 193,128 papers from the H1 Connect post-publication peer review platform. Using multilevel models, we decomposed judge-related variation into differences in overall severity and differences in the weighting of scientific attributes. We found that judge-related effects accounted for substantially more variance in ratings than the evaluated papers. In our most detailed model, judge-level effects and judge-specific slopes explained 61% of the total variance, whereas combined paper and journal-level effects accounted for only 7%. By contrast, examined measures of directional bias, such as author gender and global affiliation, explained less than 1% of the variance. We conclude that assessment outcomes were shaped more by the judges than by the papers themselves. Our results demonstrate the necessity of noise audits in high-stakes scientific evaluation.

Figures

Figures reproduced from arXiv: 2607.09783 by Alex Jones, Lutz Bornmann, Robert Ward.

Figure 1
Figure 1. Figure 1: Variance partitioning across primary models and robustness analyses. Each horizontal stacked bar shows the proportion of variance in ratings attributable to papers, journals, judge-level intercepts, judge-level slopes, and residual variation. The first three models progressively develop the full model: paper and judge intercepts; addition of journal effects; and addition of judge-specific slopes on six lat… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    H., Reisch, L

    Allison, K. H., Reisch, L. M., Carney, P. A., Weaver, D. L., Schnitt, S. J., O’Malley, F. P., & Elmore, J. G. (2014). Understanding diagnostic variability in breast pathology: Lessons learned from an expert consensus review panel. Histopathology, 65(2), 240–

  2. [2]

    Bonavia, T., & Marin-Garcia, J. A. (2023). A noise audit of the peer review of a scientific article: A WPOM journal case study. WPOM-Working Papers on Operations Management, 14(2), 137-166. https://doi.org/10.4995/wpom.19631 Bornmann, L. (2011). Scientific peer review. Annual Review of Information Science and Technology, 45, 199-245. https://doi.org/10.10...

  3. [3]

    M., Collings, A

    Murray, D., Siler, K., Larivière, V., Chan, W. M., Collings, A. M., Raymond, J., & Sugimoto, C. R. (2019). Author-reviewer homophily in peer review. Retrieved January 23, 2026, from https://www.biorxiv.org/content/biorxiv/early/2019/08/04/400515.full.pdf Pier, E. L., Brauer, M., Filut, A., Kaatz, A., Raclaw, J., Nathan, M. J., & Carnes, M. (2018). Low agr...