REVIEW 3 major objections 4 minor 66 references
At equal reviewer scores, borderline papers without a top-institution author are accepted less often at a major machine-learning conference, yet the accepted and rejected papers show no better downstream outcomes—evidence the edge comes fro
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:15 UTC pith:HC5TOQFZ
load-bearing objection The audit design is novel and the null is honest, but the headline prestige gap doesn't survive the paper's own permutation test in-sample and the out-of-sample leg is too thin to carry it. the 3 major comments →
Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that at ICLR 2019–2025, among the 10,416 borderline submissions, the equal-score acceptance gap against papers without top-25-institution authors (−0.5 percentage points in the discovery half, −1.6 in the confirmatory half) is real but does not reflect a higher bar on any measured outcome. The accepted low-prestige papers do not go on to outperform; the rejected pile does not contain systematically better papers. Instead the disparity concentrates among papers whose authors were identifiable via a pre-decision preprint (−3.4 points against −0.2) and reappears in the pre-registered 2026 cohort. The paper's stated conclusion is that the evidence is consistent with area cha
What carries the argument
Two devices carry the argument. The borderline band—submissions within half a within-year standard deviation of the fitted 50% acceptance threshold—localizes the analysis at the margin where discretion operates. The robust outcome test concludes discrimination only when a benchmark test (equal-score acceptance gap) and an outcome test (downstream quality among accepted and rejected papers) agree directionally; the paper applies it with five outcomes (citations, disruption, two novelty measures, eventual venue) and, on the reject side, traces rejected submissions to their eventual publication.
Load-bearing premise
The null assumes the five measured outcomes faithfully capture the quality an area chair targets, and that the monotone-likelihood-ratio condition holds; if prestige itself inflates these outcomes, or if the relevant quality is unmeasured, a real higher bar would leave exactly the observed null.
What would settle it
Compute the equal-score acceptance gap among borderline papers with a pre-decision preprint enrolled in an enforced anonymous-posting regime; if the −3.4-point gap persists while downstream outcomes stay balanced, the prestige-prior account fails. Conversely, in the ICLR 2026 cohort, a benchmark-concordant positive outcome disparity for low-prestige accepted or rejected papers (e.g., higher citations after tracing) would flip the null toward a higher-bar conclusion.
If this is right
- A conference's preprint policy is a fairness lever: the entire equal-score gap lives where authors are identifiable before the decision.
- Audits of discretionary gatekeeping should restrict to the margin, pre-register outcomes, and examine the rejected side when it leaves a public trace.
- The verdict 'none' means no robust-test evidence of a higher bar, not proof that the discretion is benign.
- The out-of-sample ICLR 2026 cohort recovers the decision-stage prestige gap, so the benchmark result is not a within-sample artifact.
- Because an accurate prior equalizes marginal outcomes, the outcome-test null is compatible with statistical discrimination grounded in prestige.
Where Pith is reading between the lines
- If the prestige-prior account is right, enforcing anonymous preprinting or delaying preprint visibility until after decisions should shrink or eliminate the −3.4-point gap; this is a testable policy experiment.
- The outcome null may be partly produced by the same process being tested: prestige inflates citations and venue placement, so a real higher bar could be masked by the Matthew effect; the paper's own negative citation disparities are consistent with that.
- The same borderline-band, reject-side template could be carried to journals and grant panels that publish rejection records; the paper's design suggests where to look (marginal decisions) and what to demand (concordance, not single-test significance).
- The gender axis remains under-powered and instrument-dependent; a benchmark disparity with a defensible outcome signature on fresh data would flip that verdict.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits the discretionary area-chair stage of ICLR peer review using public OpenReview records for 2019–2025 plus the pre-registered 2026 cohort. It defines a borderline band around the fitted acceptance threshold and estimates three stage-decomposition estimands: the total in-band acceptance gap, the reviewer-score gap, and the equal-score decision-stage gap (the benchmark test). It then applies a five-outcome family (citations, disruption, two novelty measures, eventual venue) to accepted and rejected borderline papers, with a pre-registered 27-cell confirmatory family and a concordance rule inspired by the Gaebler–Goel robust outcome test. The headline affirmative finding is a 0.5–1.6 percentage point lower acceptance rate for borderline papers without a top-25-institution author at equal reviewer scores, concentrated in the decision stage and, suggestively, among submissions with pre-decision arXiv preprints; the result nominally reappears on ICLR 2026. The outcome-test leg returns no benchmark-concordant evidence of a higher bar, which the paper interprets not as exoneration but as consistent with statistical discrimination through identity leakage.
Significance. If the benchmark gap were robust, this would be a valuable template for auditing discretionary gatekeeping: it brings the robust outcome-test logic to peer review, exploits the public record of rejected submissions for a reject-side outcome test, pre-registers a large confirmatory family, and reports unusually extensive robustness analysis (permutation tests, CR2 adjustments, Lee–Manski bounds, revision-extent controls, band-width and prestige-coding grids). The null outcome leg is carefully bounded and the dual-use and proxy-validity caveats are stated honestly. The main weakness is that the central affirmative claim does not survive the paper's own finite-sample permutation test in the main sample, and the out-of-sample leg is a single cohort with no multiplicity correction and no outcome data. The design and transparency are strengths, but the headline claim is currently stronger than the evidence.
major comments (3)
- [§5.1 and Appendix F] The paper's central affirmative claim — an equal-score prestige gap at the discretionary stage in ICLR 2019–2025 — does not survive its own stringent permutation test: p_perm = 0.67 (discovery) and 0.20 (confirmatory), pooled 0.27, as reported in §5.1. The text itself says the max-rule q-values 'overstate in-sample certainty.' Since this benchmark gap is the foundation of the 'identity leakage' interpretation, the abstract's '0.5 to 1.6 percentage point lower rate' and the phrase 'survive false-discovery correction' should be presented with the finite-sample failure prominently attached. Either provide an inference procedure whose conditions are met and that sustains the main-sample claim, or explicitly designate the in-sample gap as suggestive and rest the affirmative case on the out-of-sample cohort.
- [§5.5 and Table 11] The out-of-sample replication on ICLR 2026 is a single cohort with three pre-registered axes; the one significant prestige result (p_perm = 0.033) is not multiplicity-corrected. A Bonferroni correction across the three axes gives p ≈ 0.099, and a Fisher combination of the two in-sample halves plus 2026 (0.67, 0.20, 0.033) gives p ≈ 0.09. Moreover, 2026 has no outcome data, so it can support only the benchmark leg, not the robust-test conjunction. The sentence calling the 2026 result 'the strongest single evidence' therefore overstates the strength. Please pre-specify and report a multiplicity correction for the 2026 axes and temper the language accordingly.
- [§5.2 and §7] The claim that the gap 'concentrates almost entirely' among arXiv-identifiable submissions is based on a pooled-band interaction outside the confirmatory family, with no multiplicity correction, and is subject to selection into posting. The paper acknowledges these limits in §5.2, but the Discussion then elevates the pattern to a central interpretive finding ('the strongest clue,' 'nearly the whole disparity sits there'). Either move this analysis into the pre-registered family with appropriate corrections, or present it consistently as hypothesis-generating and not as part of the core evidentiary claim. As written, the abstract and Discussion give it more weight than the confirmatory status allows.
minor comments (4)
- [Abstract and §5.5] 'Reappears out-of-sample' should specify that the 2026 replication covers only the benchmark side; downstream outcomes are structurally unavailable, so the robust-test conclusion is not replicated out-of-sample.
- [§3.3 and Table 6] The prestige axis resolves for essentially no 2019–2020 submissions, so headline statements about 'ICLR 2019–2025' for prestige are effectively 2021–2025. Please state this explicitly wherever the pooled window is quoted.
- [Figure 2 and §5.2] The interaction estimates are labeled suggestive, but the figure caption should also state that the two significant interaction p-values are not corrected for the multiple axes and partitions examined.
- [§5.5] With a single venue-year cluster, the 2026 point estimate has no sandwich-based uncertainty; consider reporting a permutation or bootstrap confidence interval for β_AC in addition to the p-value.
Circularity Check
No significant circularity: the audit is pre-registered, the out-of-sample cohort is genuinely future data, and no estimand reduces by construction to a fitted input or self-citation.
full rationale
The paper's derivation chain is an observational audit against external public data (OpenReview, Semantic Scholar, CS-rankings), not a derivation that identifies its conclusion with its inputs. The central benchmark estimand β_AC is a regression-adjusted acceptance gap conditional on reviewer scores; it is not fitted to the outcome family, and the outcome family is a separately pre-registered set of downstream proxies. The ICLR 2026 result is a genuine out-of-sample replication: the cohort was collected after the plan was frozen, downstream outcomes are structurally unavailable, and the permutation test carries the inference. The robust-test verdict is explicitly conjunctive and one-directional; the paper repeatedly states that the outcome null is not proof of equal treatment and that proxies are not quality itself (Sections 4.1, 7; Appendix A). No load-bearing self-citation appears: the robust-test theorem is Gaebler and Goel [24], external work, and the paper explicitly lists its departures from that theorem, calling the design 'inspired by' rather than a direct instantiation. The acknowledged limitations (MLRP validity, proxy contestability, trace selection, few clusters, name-based gender inference) are assumptions and sensitivity analyses, not circular definitions. In particular, the 'revealed prestige prior' interpretation is presented as one reading consistent with the joint pattern, not as an inference forced by the equations. The in-sample permutation-test weakness and the fragility of the 2026 p-value are statistical-evidence concerns, not circularity. I therefore find no step in which a prediction reduces to its own input, and no self-citation chain carries the central claim.
Axiom & Free-Parameter Ledger
free parameters (4)
- borderline band half-width =
0.5 within-year SD
- top-25 prestige boundary =
CS-rankings top-25
- Q1 citation window =
3 years
- MLRP sensitivity index k, epsilon =
k=0.5, epsilon=1e-3
axioms (4)
- domain assumption MLRP holds for the conditioning signal (mean reviewer score)
- domain assumption The five outcome measures proxy the quality the AC targets
- domain assumption OpenReview public records accurately reflect review scores/decisions
- domain assumption Rejected papers' eventual published versions represent the as-rejected paper
read the original abstract
We study peer review at ICLR, a large machine-learning conference whose complete review record, including rejected submissions, is public. Reviewers score each submission; for the borderline band whose scores do not settle an outcome, an area chair makes a discretionary accept-or-reject call. We ask whether that call is even-handed: do authors from prestigious institutions, WEIRD countries, or all-male teams get the benefit of the doubt at the margin? Across ICLR 2019-2025 (31,711 submissions; 10,416 borderline), borderline papers without a top-25-institution author are accepted at a 0.5 to 1.6 percentage point lower rate at the same reviewer scores. The gap arises at the discretionary stage, reappears out-of-sample in the pre-registered ICLR 2026 cohort, and concentrates almost entirely among submissions identifiable through a pre-decision arXiv preprint (-3.4 vs. -0.2 points). Equal scores need not mean equal papers: an area chair may respond to quality the scores miss. We apply a robust outcome test, which concludes discrimination only when the group accepted at a lower rate also realizes better downstream outcomes. We measure five outcomes (citations, disruption, two forms of novelty, eventual venue) on both sides of the decision, including the first "ones that got away" test of rejected submissions. Our headline result is a null: across a pre-registered family of 27 tests, no disparity concordant with the decision-rate gap survives correction; we find no evidence that any group faced a higher bar on the outcomes we measure. That null is not an exoneration. A pre-decision preprint pierces the blind through policy-permitted means, and the acceptance gap lives almost entirely in that porosity, consistent with area chairs using revealed institutional prestige as a prior: statistical discrimination that outcome tests may not detect, and a practice double-blind review exists to prevent.
Figures
Reference graph
Works this paper leans on
-
[1]
Aksnes, Liv Langfeldt, and Paul Wouters
Dag W. Aksnes, Liv Langfeldt, and Paul Wouters. 2019. Citations, citation indicators, and research quality: An overview of basic concepts and theories.SAGE Open9, 1 (2019). https://doi.org/10.1177/2158244019829575
-
[2]
Shamena Anwar and Hanming Fang. 2006. An alternative test of racial prejudice in motor vehicle searches: Theory and evidence.American Economic Review96, 1 (2006), 127–151. https://doi.org/10.1257/000282806776157579
-
[3]
David Arnold, Will Dobbie, and Crystal S. Yang. 2018. Racial bias in bail decisions.The Quarterly Journal of Economics133, 4 (2018), 1885–1932. https://doi.org/10.1093/qje/qjy012 16 Hazem Ibrahim, Talal Rahwan, and Yasir Zaki
-
[4]
Kenneth J. Arrow. 1973. The theory of discrimination. InDiscrimination in Labor Markets, Orley Ashenfelter and Albert Rees (Eds.). Princeton University Press, Princeton, NJ, 3–33
1973
-
[5]
Ian Ayres. 2002. Outcome tests of racial disparities in police practices.Justice Research and Policy4, 1–2 (2002), 131–142. https://doi.org/10.3818/ JRP.4.1.2002.131
2002
-
[6]
Gary S. Becker. 1971.The Economics of Discrimination(2 ed.). University of Chicago Press, Chicago, IL
1971
-
[7]
Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B (Methodological)57, 1 (1995), 289–300
1995
-
[8]
Marianne Bertrand and Sendhil Mullainathan. 2004. Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination.American Economic Review94, 4 (2004), 991–1013. https://doi.org/10.1257/0002828042002561
-
[9]
Dauphin, Percy Liang, and Jennifer Wortman Vaughan
Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021. The NeurIPS 2021 consistency experiment. NeurIPS Blog. Fuller write-up: arXiv:2306.03262
Pith/arXiv arXiv 2021
-
[10]
Homanga Bharadhwaj, Dylan Turpin, Animesh Garg, and Ashton Anderson. 2020. De-anonymization of authors through arXiv submissions during double-blind review.arXiv preprint arXiv:2007.00177(2020)
Pith/arXiv arXiv 2020
-
[11]
Rebecca M. Blank. 1991. The effects of double-blind versus single-blind reviewing: Experimental evidence from The American Economic Review. American Economic Review81, 5 (1991), 1041–1067
1991
-
[12]
Aislinn Bohren, Alex Imas, and Michael Rosenberg
J. Aislinn Bohren, Alex Imas, and Michael Rosenberg. 2019. The dynamics of discrimination: Theory and evidence.American Economic Review109, 10 (2019), 3395–3436. https://doi.org/10.1257/aer.20171829
-
[13]
Lutz Bornmann. 2011. Scientific peer review.Annual Review of Information Science and Technology45, 1 (2011), 197–245. https://doi.org/10.1002/ aris.2011.1440450112
Pith/arXiv arXiv 2011
-
[14]
Lutz Bornmann and Hans-Dieter Daniel. 2008. What do citation counts measure? A review of studies on citing behavior.Journal of Documentation 64, 1 (2008), 45–80. https://doi.org/10.1108/00220410810844150
-
[15]
Kevin J. Boudreau, Eva C. Guinan, Karim R. Lakhani, and Christoph Riedl. 2016. Looking across and looking beyond the knowledge frontier: Intellectual distance, novelty, and resource allocation in science.Management Science62, 10 (2016), 2765–2783. https://doi.org/10.1287/mnsc.2015.2285
arXiv 2016
-
[16]
Budden, Tom Tregenza, Lonnie W
Amber E. Budden, Tom Tregenza, Lonnie W. Aarssen, Julia Koricheva, Roosa Leimu, and Christopher J. Lortie. 2008. Double-blind review favours increased representation of female authors.Trends in Ecology & Evolution23, 1 (2008), 4–6. https://doi.org/10.1016/j.tree.2007.07.008
-
[17]
Canay, Magne Mogstad, and Jack Mountjoy
Ivan A. Canay, Magne Mogstad, and Jack Mountjoy. 2024. On the use of outcome tests for detecting bias in decision making.The Review of Economic Studies91, 4 (2024), 2135–2167. https://doi.org/10.1093/restud/rdad082
-
[18]
David Card, Stefano DellaVigna, Patricia Funk, and Nagore Iriberri. 2020. Are referees and editors in economics gender neutral?The Quarterly Journal of Economics135, 1 (2020), 269–327. https://doi.org/10.1093/qje/qjz035
-
[19]
Aaron Clauset, Samuel Arbesman, and Daniel B. Larremore. 2015. Systematic inequality and hierarchy in faculty hiring networks.Science Advances 1, 1 (2015), e1400005. https://doi.org/10.1126/sciadv.1400005
-
[20]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. SPECTER: Document-level representation learning using citation-informed transformers. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2270–2282
2020
-
[21]
David Firth. 1993. Bias reduction of maximum likelihood estimates.Biometrika80, 1 (1993), 27–38. https://doi.org/10.1093/biomet/80.1.27
-
[22]
Bergstrom, Katy Börner, James A
Santo Fortunato, Carl T. Bergstrom, Katy Börner, James A. Evans, Dirk Helbing, Staša Milojévić, Alexander M. Petersen, Filippo Radicchi, Roberta Sinatra, Brian Uzzi, Alessandro Vespignani, Ludo Waltman, Dashun Wang, and Albert-László Barabási. 2018. Science of science.Science359, 6379 (2018), eaao0185. https://doi.org/10.1126/science.aao0185
-
[23]
Funk and Jason Owen-Smith
Russell J. Funk and Jason Owen-Smith. 2017. A dynamic network measure of technological change.Management Science63, 3 (2017), 791–817
2017
-
[24]
Johann D. Gaebler and Sharad Goel. 2025. A simple, statistically robust test of discrimination.Proceedings of the National Academy of Sciences122, 10 (2025), e2416348122. https://doi.org/10.1073/pnas.2416348122
-
[25]
Donna K. Ginther, Walter T. Schaffer, Joshua Schnell, Beth Masimore, Faye Liu, Laurel L. Haak, and Raynard Kington. 2011. Race, ethnicity, and NIH research awards.Science333, 6045 (2011), 1015–1019. https://doi.org/10.1126/science.1196783
-
[26]
Joseph Henrich, Steven J. Heine, and Ara Norenzayan. 2010. The weirdest people in the world?Behavioral and Brain Sciences33, 2-3 (2010), 61–83. https://doi.org/10.1017/S0140525X0999152X
-
[27]
Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A
Bas Hofstra, Vivek V. Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A. McFarland. 2020. The diversity–innovation paradox in science.Proceedings of the National Academy of Sciences117, 17 (2020), 9284–9291. https://doi.org/10.1073/pnas.1915378117
-
[28]
D. G. Horvitz and D. J. Thompson. 1952. A generalization of sampling without replacement from a finite universe.J. Amer. Statist. Assoc.47, 260 (1952), 663–685. https://doi.org/10.1080/01621459.1952.10483446
arXiv 1952
-
[29]
Jürgen Huber, Sabiou Inoua, Rudolf Kerschbamer, Christian König-Kersting, Stefan Palan, and Vernon L. Smith. 2022. Nobel and novice: Author prominence affects peer review.Proceedings of the National Academy of Sciences119, 41 (2022), e2205779119. https://doi.org/10.1073/pnas.2205779119
-
[30]
2021.What marginal outcome tests can tell us about racially biased decision-making
Peter Hull. 2021.What marginal outcome tests can tell us about racially biased decision-making. Working Paper 28503. National Bureau of Economic Research. https://doi.org/10.3386/w28503
-
[31]
Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, et al. 2023. The Semantic Scholar open data platform.arXiv preprint arXiv:2301.10140(2023)
Pith/arXiv arXiv 2023
-
[32]
Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan. 2018. Human decisions and machine predictions.The Quarterly Journal of Economics133, 1 (2018), 237–293. https://doi.org/10.1093/qje/qjx032 Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review? Evidence from ICLR 17
-
[33]
John Knowles, Nicola Persico, and Petra Todd. 2001. Racial bias in motor vehicle searches: Theory and evidence.Journal of Political Economy109, 1 (2001), 203–229. https://doi.org/10.1086/318603
-
[34]
Daniël Lakens. 2017. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses.Social Psychological and Personality Science8, 4 (2017), 355–362. https://doi.org/10.1177/1948550617697177
-
[35]
John Langford and Mark Guzdial. 2015. The arbitrariness of reviews, and advice for school administrators.Commun. ACM58, 4 (2015), 12–13. https://doi.org/10.1145/2732417
doi:10.1145/2732417 2015
-
[36]
Carole J. Lee, Cassidy R. Sugimoto, Guo Zhang, and Blaise Cronin. 2013. Bias in peer review.Journal of the American Society for Information Science and Technology64, 1 (2013), 2–17. https://doi.org/10.1002/asi.22784
-
[37]
David S. Lee. 2009. Training, wages, and sample selection: Estimating sharp bounds on treatment effects.The Review of Economic Studies76, 3 (2009), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x
Pith/arXiv arXiv 2009
-
[38]
David S. Lee and Thomas Lemieux. 2010. Regression discontinuity designs in economics.Journal of Economic Literature48, 2 (2010), 281–355. https://doi.org/10.1257/jel.48.2.281
-
[39]
Haley Lepp and Daniel Scott Smith. 2025. “You Cannot Sound Like GPT”: Signs of language discrimination and resistance in computer science publishing. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 3162–3181. https://doi.org/10.1145/3715275. 3732202
doi:10.1145/3715275 2025
-
[40]
Charles F. Manski. 1990. Nonparametric bounds on treatment effects.The American Economic Review80, 2 (1990), 319–323. Papers and Proceedings
1990
-
[41]
Emaad Manzoor and Nihar B. Shah. 2021. Uncovering latent biases in text: Method and application to peer review. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 4767–4775. https://doi.org/10.1609/aaai.v35i6.16608
-
[42]
Robert K. Merton. 1968. The Matthew effect in science.Science159, 3810 (1968), 56–63. https://doi.org/10.1126/science.159.3810.56
-
[43]
Kanu Okike, Kevin T. Hug, Mininder S. Kocher, and Seth S. Leopold. 2016. Single-blind vs double-blind peer review in the setting of author prestige. JAMA316, 12 (2016), 1315–1316. https://doi.org/10.1001/jama.2016.11014
arXiv 2016
-
[44]
Michael Park, Erin Leahey, and Russell J. Funk. 2023. Papers and patents are becoming less disruptive over time.Nature613 (2023), 138–144. https://doi.org/10.1038/s41586-022-05543-x
-
[45]
Douglas P. Peters and Stephen J. Ceci. 1982. Peer-review practices of psychological journals: The fate of published articles, submitted again. Behavioral and Brain Sciences5, 2 (1982), 187–195. https://doi.org/10.1017/S0140525X00011183
-
[46]
Edmund S. Phelps. 1972. The statistical theory of racism and sexism.American Economic Review62, 4 (1972), 659–661
1972
-
[47]
Pier, Markus Brauer, Amarette Filut, Anna Kaatz, Joshua Raclaw, Mitchell J
Elizabeth L. Pier, Markus Brauer, Amarette Filut, Anna Kaatz, Joshua Raclaw, Mitchell J. Nathan, Cecilia E. Ford, and Molly Carnes. 2018. Low agreement among reviewers evaluating the same NIH grant applications.Proceedings of the National Academy of Sciences115, 12 (2018), 2952–2957. https://doi.org/10.1073/pnas.1714379115
-
[48]
Emma Pierson, Camelia Simoiu, Jan Overgoor, Sam Corbett-Davies, Daniel Jenson, Amy Shoemaker, Vignesh Ramachandran, Phoebe Barghouty, Cheryl Phillips, Ravi Shroff, and Sharad Goel. 2020. A large-scale analysis of racial disparities in police stops across the United States.Nature Human Behaviour4, 7 (2020), 736–745. https://doi.org/10.1038/s41562-020-0858-1
-
[49]
Charvi Rastogi, Ivan Stelmakh, Xinwei Shen, Marina Meila, Federico Echenique, Shuchi Chawla, and Nihar B. Shah. 2022. To arXiv or not to arXiv: A study quantifying pros and cons of posting preprints online. arXiv:2203.17259
arXiv 2022
-
[50]
Joseph S. Ross, Cary P. Gross, Mayur M. Desai, Yuling Hong, Augustus O. Grant, Stephen R. Daniels, Vladimir C. Hachinski, Raymond J. Gibbons, Timothy J. Gardner, and Harlan M. Krumholz. 2006. Effect of blinded peer review on abstract acceptance.JAMA295, 14 (2006), 1675–1680. https://doi.org/10.1001/jama.295.14.1675
-
[51]
Davidson, Veniamin Veselovsky, and Robert West
Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, and Robert West. 2025. The AI review lottery: Widespread AI-assisted peer reviews boost paper scores and acceptance rates.Proceedings of the ACM on Human-Computer Interaction9, 7 (CSCW) (2025). https://doi.org/10.1145/3757667
doi:10.1145/3757667 2025
-
[52]
Lucía Santamaría and Helena Mihaljević. 2018. Comparison and benchmark of name-to-gender inference services.PeerJ Computer Science4 (2018), e156. https://doi.org/10.7717/peerj-cs.156
-
[53]
Donald J. Schuirmann. 1987. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability.Journal of Pharmacokinetics and Biopharmaceutics15, 6 (1987), 657–680. https://doi.org/10.1007/BF01068419
-
[54]
Nihar B. Shah. 2022. Challenges, experiments, and computational solutions in peer review.Commun. ACM65, 6 (2022), 76–87. https://doi.org/10. 1145/3528086
2022
-
[55]
Camelia Simoiu, Sam Corbett-Davies, and Sharad Goel. 2017. The problem of infra-marginality in outcome tests for discrimination.Annals of Applied Statistics11, 3 (2017), 1193–1216. https://doi.org/10.1214/17-AOAS1058
-
[56]
Flaminio Squazzoni, Giangiacomo Bravo, Mike Farjam, Ana Marušić, Bahar Mehmani, Michael Willis, Aliaksandr Birukou, Pierpaolo Dondio, and Francisco Grimaldo. 2021. Peer review and gender bias: A study on 145 scholarly journals.Science Advances7, 2 (2021), eabd0299. https: //doi.org/10.1126/sciadv.abd0299
-
[57]
Ivan Stelmakh, Charvi Rastogi, Ryan Liu, Shuchi Chawla, Federico Echenique, and Nihar B. Shah. 2023. Cite-seeing and reviewing: A study on citation bias in peer review.PLOS ONE18, 7 (2023), e0283980. https://doi.org/10.1371/journal.pone.0283980
-
[58]
Shah, Aarti Singh, and Hal Daumé III
Ivan Stelmakh, Nihar B. Shah, Aarti Singh, and Hal Daumé III. 2021. Prior and prejudice: The novice reviewers’ bias against resubmissions in conference peer review.Proceedings of the ACM on Human-Computer Interaction5, CSCW1, Article 75 (2021). https://doi.org/10.1145/3449149 18 Hazem Ibrahim, Talal Rahwan, and Yasir Zaki
doi:10.1145/3449149 2021
-
[59]
Misha Teplitskiy, Daniel Acuna, Aïda Elamrani-Raoult, Konrad Körding, and James Evans. 2018. The sociology of scientific validity: How professional networks shape judgement in peer review.Research Policy47, 9 (2018), 1825–1841. https://doi.org/10.1016/j.respol.2018.06.014
-
[60]
Andrew Tomkins, Min Zhang, and William D. Heavlin. 2017. Reviewer bias in single- versus double-blind peer review.Proceedings of the National Academy of Sciences114, 48 (2017), 12708–12713. https://doi.org/10.1073/pnas.1707323114
-
[61]
David Tran, Alex Valtchanov, Keshav Ganapathy, Raymond Feng, Eric Slud, Micah Goldblum, and Tom Goldstein. 2020. An open review of OpenReview: A critical analysis of the machine learning conference review process. InNeurIPS 2020 Workshop on Navigating the Broader Impacts of AI Research. Workshop paper, arXiv:2010.05137
Pith/arXiv arXiv 2020
-
[62]
Brian Uzzi, Satyam Mukherjee, Michael Stringer, and Ben Jones. 2013. Atypical combinations and scientific impact.Science342, 6157 (2013), 468–472
2013
-
[63]
Jian Wang, Reinhilde Veugelers, and Paula Stephan. 2017. Bias against novelty in science: A cautionary tale for users of bibliometric indicators. Research Policy46, 8 (2017), 1416–1436. https://doi.org/10.1016/j.respol.2017.06.006
-
[64]
Samuel F. Way, Allison C. Morgan, Daniel B. Larremore, and Aaron Clauset. 2019. Productivity, prominence, and the effects of academic environment. Proceedings of the National Academy of Sciences116, 22 (2019), 10729–10733. https://doi.org/10.1073/pnas.1817431116
-
[65]
Witteman, Michael Hendricks, Sharon Straus, and Cara Tannenbaum
Holly O. Witteman, Michael Hendricks, Sharon Straus, and Cara Tannenbaum. 2019. Are gender gaps due to evaluations of the applicant or the science? A natural experiment at a national funding agency.The Lancet393, 10171 (2019), 531–540. https://doi.org/10.1016/S0140-6736(18)32611-4
-
[66]
Lingfei Wu, Dashun Wang, and James A. Evans. 2019. Large teams develop and small teams disrupt science and technology.Nature566 (2019), 378–382. A Relation to the robust-test theorem Our implementation departs from the canonical form of the Gaebler–Goel result in three ways, so we present the design as inspired by the robust test rather than a direct inst...
arXiv 2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.