Pith. sign in

REVIEW 8 cited by

Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.09774 v1 pith:ZD6SBDBV submitted 2021-09-20 cs.DL cs.LG

classification cs.DLcs.LG
keywords experimentconferencequalityneuripsscorescorrelationfindgood
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we revisit the 2014 NeurIPS experiment that examined inconsistency in conference peer review. We determine that 50\% of the variation in reviewer quality scores was subjective in origin. Further, with seven years passing since the experiment we find that for \emph{accepted} papers, there is no correlation between quality scores and impact of the paper as measured as a function of citation count. We trace the fate of rejected papers, recovering where these papers were eventually published. For these papers we find a correlation between quality scores and impact. We conclude that the reviewing process for the 2014 conference was good for identifying poor papers, but poor for identifying good papers. We give some suggestions for improving the reviewing process but also warn against removing the subjective element. Finally, we suggest that the real conclusion of the experiment is that the community should place less onus on the notion of `top-tier conference publications' when assessing the quality of individual researchers. For NeurIPS 2021, the PCs are repeating the experiment, as well as conducting new ones.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can AI agents conduct open-ended AI research? Early evidence from two case studies

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Shadow evaluations on two unpublished NeurIPS questions show today's agents can engineer AI experiments but cannot produce publishable open-ended research.

  2. Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025

    cs.DL 2026-05 conditional novelty 7.0 of 10

    ICLR peer-review scores are orthogonal to future disruptiveness (EDM); EDM identifies highly cited catalysts far better than CD, node2vec, or an LLM rater, and catalyst types precede large topic-share and cross-topic ...

  3. ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review

    cs.HC 2026-07 conditional novelty 6.0 of 10

    ReVoicer is a prototype that turns spoken, in-the-moment reactions to a paper into cleaned, tagged annotations and a draft review aligned with the reviewer's own style.

  4. Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

    cs.DL 2026-04 conditional novelty 6.0 of 10

    An audit of 50,289 ICLR papers shows acceptance odds vary up to 8x across topics at equal reviewer scores, indicating scores are not comparable across research areas.

  5. Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

    cs.LG 2025-06 conditional novelty 5.0 of 10

    ML conferences should create an official peer-reviewed track dedicated to refuting and critiquing previously published work.

  6. A Vision for the Future of an AI-Integrated Research Ecosystem

    cs.CY 2026-08 unverdicted novelty 4.0 of 10

    Scientific publishing should shift from detecting AI-generated content to building infrastructure for provenance, calibration, and accountability, making trustworthy AI-assisted scholarship the default.

  7. Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper argues that AI-assisted peer review is an urgent priority and that its success depends on collecting richer, structured peer review process data.

  8. AI Scientists Fail Without Strong Implementation Capability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    AI scientist systems can propose ideas but cannot reliably implement and verify experiments, making the implementation gap, not idea generation, the current bottleneck.

Pith tools