REVIEW 4 cited by
An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mainstream machine learning conferences have seen a dramatic increase in the number of participants, along with a growing range of perspectives, in recent years. Members of the machine learning community are likely to overhear allegations ranging from randomness of acceptance decisions to institutional bias. In this work, we critically analyze the review process through a comprehensive study of papers submitted to ICLR between 2017 and 2020. We quantify reproducibility/randomness in review scores and acceptance decisions, and examine whether scores correlate with paper impact. Our findings suggest strong institutional bias in accept/reject decisions, even after controlling for paper quality. Furthermore, we find evidence for a gender gap, with female authors receiving lower scores, lower acceptance rates, and fewer citations per paper than their male counterparts. We conclude our work with recommendations for future conference organizers.
Forward citations
Cited by 4 Pith papers
-
Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?
At ICLR, equally scored borderline papers from outside top-25 institutions are accepted less often, a gap concentrated in preprint-identifiable submissions; outcome tests find no evidence of a higher bar.
-
Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025
ICLR peer-review scores are orthogonal to future disruptiveness (EDM); EDM identifies highly cited catalysts far better than CD, node2vec, or an LLM rater, and catalyst types precede large topic-share and cross-topic ...
-
Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
An audit of 50,289 ICLR papers shows acceptance odds vary up to 8x across topics at equal reviewer scores, indicating scores are not comparable across research areas.
-
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.
Discussion (0). Continue with ORCID to comment.