Pith. sign in

REVIEW 4 major objections 5 minor 52 references

Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Assigning reviewers who differ in co-authorship, seniority, or topic changes what the collective reviews cover and how much they repeat, while geographical diversity does neither.

desk verdict A real new empirical question with a useful null result, but the within-paper differencing doesn't cancel the anchor reviewer and the effect sizes are not credible as written. read the letter →

arxiv 2411.11437 v1 pith:KPHGRLT2 submitted 2024-11-18 cs.DL cs.CLstat.AP

classification cs.DLcs.CLstat.AP
keywords peerreviewreviewerassignmentdiversitycoverageredundancycausalinferenceobservationalstudygroupdecision-making
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the mix of reviewers assigned to a paper changes what the reviews, taken together, cover and whether they repeat one another. Using roughly 5,000 submissions from a machine-learning conference, it treats five axes of reviewer diversity—organization, geography, seniority, topic, and co-authorship—as treatments and measures two group-level review outcomes: coverage (of the paper's content and of review criteria) and redundancy (overlap between reviews). The central finding is that the axes act differently: co-authorship and seniority diversity raise coverage of review types, topical diversity raises coverage of the paper's content, and all axes except geography reduce redundancy. If the causal reading holds, program chairs can nudge review utility by adding specific diversity constraints to assignment algorithms without sacrificing reviewer expertise.

What carries the argument

The load-bearing device is within-paper pair differencing. Each submission is reviewed by three or four reviewers, so the analysis builds triples $(r_1, r_2, r_3)$ in which $(r_1, r_2)$ is diverse and $(r_1, r_3)$ is non-diverse along a target dimension $d^*$. Subtracting the two linear outcome equations removes the submission term $\omega_S$, so paper content and quality drop out of the comparison; the remaining regression estimates $\gamma_{d^*}$ with controls for the other diversity dimensions, each reviewer's profile vector, and the TPMS expertise scores $E(r, S)$. The same logic is applied non-parametrically by propensity-score matching diverse and non-diverse pairs within a paper.

What would settle it

A randomized assignment experiment would settle the claim: for the same set of papers, randomly assign reviewer pairs that are matched on TPMS expertise and profile but differ only in co-authorship or topical diversity, then compare coverage and redundancy; if the differences vanish, the observational effects are confounded. Short of that, a sensitivity analysis introducing an unmeasured confounder (for example, review length or time spent) into Equation 3 would reveal whether the estimated $\gamma$ values survive plausible confounding.

Watch

Extended reading notes

Core claim

The paper's central claim is that reviewer-slate diversity has causal, dimension-specific effects on review coverage and redundancy. In the parametric analysis, co-authorship diversity increases argument- and aspect-type coverage (effects 0.0064 and 0.0059) and seniority diversity increases the same (0.0058 and 0.0074), while topical diversity increases lexical and semantic paper coverage (0.1098 and 0.0929). Lexical and semantic redundancy decrease under organizational (−0.0218, −0.0078), co-authorship (−0.0258, −0.0097), topical (−0.0524, −0.0290), and seniority (−0.0061, −0.0015) diversity; geography shows no significant effect on any outcome. Co-authorship diversity is the only axis that also lowers redundancy within a single review criterion (weighted semantic redundancy). The paper reports that non-parametric propensity-score matching corroborates the significant parametric results.

Load-bearing premise

After differencing within a paper and controlling for reviewer profiles and TPMS expertise, the paper assumes the only systematic reason diverse and non-diverse pairs differ in coverage or redundancy is the diversity itself—an untested assumption that review length, review effort, or hidden assignment constraints are not also driving the differences.

Editorial extensions

If this is right

  • Reviewer assignment systems can target co-authorship diversity (no shared co-authors) to increase type coverage and lower redundancy, including within individual review criteria.
  • Topical diversity is the lever for covering more of the paper's content itself, while it does not raise coverage of review criteria.
  • Organizational and seniority diversity reduce redundancy but do not by themselves raise paper coverage, so the choice of diversity axis depends on the desired outcome.
  • Geographical diversity shows no effect on coverage or redundancy, so geography-based assignment constraints cannot be justified by review-utility gains.
  • The effects are not explained by lower review quality: diversity shows no significant correlation with meta-reviewers' ratings of review quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The within-paper differencing design transfers to any conference dataset with three or more reviewers per paper, reviewer profiles, and assignment-expertise scores; the causal contrast does not require outcome data from re-reviewing the same paper.
  • If geographical diversity truly has no effect on coverage or redundancy, venues that add geography constraints to assignments for other reasons (for example, reducing collusion) should treat those constraints as orthogonal to review breadth—the paper's evidence does not speak to collusion outcomes.
  • A useful next experiment would separate 'co-authors' from 'shared co-author' pairs: the paper's co-authorship diversity collapses both into the non-diverse category, so the active mechanism could be direct collaboration rather than network proximity.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies whether reviewer-slate diversity along five axes (organization, geography, seniority, topic, and co-authorship) causally affects two group-level review outcomes: coverage of the paper and review criteria, and redundancy across reviews. Using ICML 2020 data with nearly 5,000 submissions, the authors propose lexical, semantic, and type-based outcome measures, then estimate effects with a within-paper differencing regression and, as robustness, a propensity-score-matched permutation analysis. They report that co-authorship and seniority diversity increase type coverage, topical diversity increases paper coverage, and organizational, seniority, topical, and co-authorship diversity reduce redundancy, while geographical diversity has no significant effect. The paper concludes with policy recommendations for reviewer assignment systems. The central claim is causal: that the estimated diversity coefficients in Table 1 can guide assignment policies to improve review utility.

Significance. If the estimates were valid, the work would be a valuable and actionable contribution. It operationalizes two understudied group-level review properties, examines five diversity axes on a large real conference dataset, includes human validation for some semantic measures, applies multiple-testing correction, and provides a non-parametric robustness check. The paper is also unusually clear in translating its findings into concrete assignment-policy recommendations. However, the main estimand is compromised by a non-additivity problem in the outcome measures, and the reported numerical results appear to contain a factor-of-two error; these issues affect the central causal claims and require a substantive reanalysis before the conclusions can be relied upon.

major comments (4)
  1. [Section 5.5, Eq. (3); Appendix A, Eqs. (4)-(8)] The anchor-cancellation argument is invalid for the outcome definitions used in the paper. In Eq. (3) the paper subtracts y(r1,r2;S) - y(r1,r3;S) and claims that the common reviewer r1 does not contribute to the difference. That is true for the linear profile and expertise covariates, but the outcome y itself is not additively separable in R1 and R2: lexical coverage (Eq. 4) uses the union of n-grams over the concatenated reviews, semantic coverage (Eq. 5) takes a max over the concatenated review set, lexical redundancy (Eq. 6) is the intersection of R1 and R2, and semantic redundancy (Eqs. 7-8) involves max or all-pair similarities across the two texts. Thus R1 still enters the difference through union, intersection, max, and pairwise product operations. For example, if the anchor review already covers every abstract n-gram, the lexical-coverage difference between a diverse and a non-diverse pair is zero regardless of the second reviewer; if the anchor has no overlap, the same difference is driven entirely by r2 versus r3. Because anchors are not randomly chosen and anchor writing style is not controlled, the gamma estimates in Table 1 can reflect an interaction between anchor style and pair composition rather than a causal effect of diversity. The non-parametric difference in Appendix D inherits the same problem.
  2. [Section 5.5, Eqs. (2a), (2b), and (3)] The algebra of the within-paper difference drops a factor of two. With delta_d*(r1,r2)=1 and delta_d*(r1,r3)=-1, subtracting Eq. (2b) from Eq. (2a) gives y(r1,r2;S) - y(r1,r3;S) approximately equal to 2*gamma_d* plus the covariate-difference terms, not gamma_d* as written in Eq. (3). Unless the model is reparameterized so that the coefficient in Eq. (3) is gamma_d*/2, the estimates reported in Table 1 are twice the gamma_d* defined in Eq. (1). The authors should state which convention was used and correct either the equations or the reported effect sizes.
  3. [Appendix C; Table 1; Section 6.2] The human validation undermines one of the outcome measures used as evidence. Appendix C reports that lexical redundancy agreed with human annotators only at chance level (p=0.623), yet Table 1 and Section 6.2 treat significant lexical-redundancy effects for organizational, seniority, topical, and co-authorship diversity as support for Hypothesis 2. In addition, the coverage measures used for Hypothesis 1 are not included in the human annotation study at all, despite Section 5.2 saying that the automated outcome measures were validated. The paper should either restrict its evidentiary claims to the human-validated semantic measures or provide separate validation for lexical redundancy and the coverage measures before using them as the basis for policy recommendations.
  4. [Section 5.5; Figure 1; Section 5.3] The causal reading of gamma_d* requires conditional ignorability after controlling for submission, reviewer profiles, and TPMS expertise, but this assumption is not probed. Review length, review effort, and the optimization constraints of the actual assignment algorithm are plausible unmeasured confounders that could correlate with both pair diversity and coverage/redundancy, and no sensitivity analysis addresses them. In addition, the treatment of missing profile data (51.0% missing Semantic Scholar profiles and 17.5% missing Google Scholar profiles) as zeros is asserted in Section 5.4 without a missingness analysis or complete-case robustness check. These gaps should be addressed before the results can support the paper's causal language and policy conclusions.
minor comments (5)
  1. [Table 1] Table 1 reports only point estimates and significance stars; standard errors or confidence intervals should be included, especially because many significant effect sizes are small in magnitude.
  2. [Section 5.5] The phrase 'without loss of generality' is not accurate for selecting an anchor reviewer and a diverse/non-diverse pair among papers with three or four reviews; the paper should specify how anchors and pairs are selected and whether estimates are averaged over all valid pairings.
  3. [Section 5.4] The co-authorship diversity measure should define 'co-authorship distance' explicitly, for example as the shortest path in the co-author graph, and the treatment of missing profile vectors as delta=0 should be stated more precisely than 'ensures a consistent estimation.'
  4. [Appendix A/B] The normalization to [0,1] described in Appendix B is not shown as a formula; provide the exact transformation for each outcome, especially for the non-symmetric weighted semantic redundancy in Eq. (8).
  5. [Section 5.2.1] For type coverage, the paper should state whether the DistilBERT aspect and argument classifiers were applied to individual review sentences or to the concatenated review text, and how the reported 90.11% and 82.96% accuracies were measured.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: diversity treatments and coverage/redundancy outcomes are measured from independent sources, outcome normalization uses external ICLR data, and the causal estimates are not fitted predictions or definitional identities.

full rationale

The paper's derivation chain is self-contained and does not reduce to its own inputs. Treatments are defined from reviewer profile fields and TPMS expertise scores; outcomes are computed from review text and paper abstracts using classifier and embedding models, with normalization ranges estimated from a separate ICLR 2019 dataset to avoid contamination. The within-paper differencing in Eq. 3 estimates gamma_{d*} as a regression coefficient rather than renaming a fitted value as a prediction. The self-citations (e.g., Stelmakh et al. 2023a for linear modeling, Shah 2022 for diversity motivation) are not load-bearing: the linear model is a standard convention and the diversity hypotheses are not justified by a uniqueness theorem or by an unverified prior result from the same authors. The Limitations section acknowledges scope restrictions and the pre-rebuttal focus, but does not assert any circular dependency. Although the anchor-cancellation step in Section 5.5 rests on an additivity assumption that may be questionable given the non-additive outcome formulas in Appendix A, that is a statistical validity concern rather than circularity: the outcome y(r1,r2;S) is not defined in terms of the treatment delta_d, and Eq. 3 does not make the estimate equivalent to its input by construction. Accordingly, no circular step can be exhibited, and the paper receives a score of 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The causal estimates rest on assumptions of conditional ignorability, linearity, no interference, and ignorable missingness. The treatment thresholds are data-derived, making the definition of 'diverse' sample-dependent. No new entities are postulated.

free parameters (5)
  • Topical diversity binarization threshold = 0.6472 (median topical similarity)
    Dot-product similarity of LDA topic vectors is binarized at the median in Section 5.4; the treatment definition is sample-dependent, though not fitted to the outcome.
  • LDA topic count = 10
    The number of topics is chosen by optimizing topic coherence on the same reviewer abstracts (Appendix E), which affects the topical diversity measure.
  • Seniority median h-index = 22
    Reviewers are split into senior and non-senior at the median h-index from Google Scholar profiles (Section 5.3); the threshold is data-derived.
  • Co-authorship distance threshold = 2
    Reviewers with co-authorship distance at most 2 are considered non-diverse (Section 5.4); the threshold determines the treatment definition.
  • Propensity score caliper = 0.1
    Matching tolerance used in the non-parametric approach (Appendix D); it affects the matched sample sizes.
assumptions (5)
  • domain assumption Conditional ignorability: no unmeasured confounders after controlling for paper, reviewer profile, and TPMS expertise.
    Needed to interpret the estimated coefficients as causal effects; not tested with sensitivity analysis (Section 5.5, Figure 1).
  • domain assumption Linearity of the outcome model in Equation 1.
    The parametric approach assumes linear additive effects of diversity, profiles, and expertise; Equation 3 is derived from this linear model.
  • domain assumption No interference between reviewers within a paper (SUTVA).
    The analysis uses pre-rebuttal reviews to avoid discussion leakage, but pair outcomes may still depend on which other reviewers are assigned to the paper (Section 5.1).
  • ad hoc to paper Missing profile data are ignorable and zero-imputation does not bias estimates.
    Profile information is missing for up to 51% of reviewers on Semantic Scholar and 17.5% on Google Scholar; missing vectors are set to 0 and diversity to 0 (Section 5.4), with no missingness model.
  • domain assumption Sentence-BERT cosine similarity and classifier labels capture review coverage and redundancy.
    Human validation is provided only for semantic and weighted semantic redundancy; lexical redundancy agreement was random and coverage measures were not human-validated (Appendix C).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing." pith.science (2026). https://pith.science/paper/KPHGRLT2

@misc{pith2026241111437,
  author       = {Pith},
  title        = {Pith review of: Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KPHGRLT2}},
  note         = {Machine review of arXiv:2411.11437}
}
read the original abstract

A large host of scientific journals and conferences solicit peer reviews from multiple reviewers for the same submission, aiming to gather a broader range of perspectives and mitigate individual biases. In this work, we reflect on the role of diversity in the slate of reviewers assigned to evaluate a submitted paper as a factor in diversifying perspectives and improving the utility of the peer-review process. We propose two measures for assessing review utility: review coverage -- reviews should cover most contents of the paper -- and review redundancy -- reviews should add information not already present in other reviews. We hypothesize that reviews from diverse reviewers will exhibit high coverage and low redundancy. We conduct a causal study of different measures of reviewer diversity on review coverage and redundancy using observational data from a peer-reviewed conference with approximately 5,000 submitted papers. Our study reveals disparate effects of different diversity measures on review coverage and redundancy. Our study finds that assigning a group of reviewers that are topically diverse, have different seniority levels, or have distinct publication networks leads to broader coverage of the paper or review criteria, but we find no evidence of an increase in coverage for reviewer slates with reviewers from diverse organizations or geographical locations. Reviewers from different organizations, seniority levels, topics, or publications networks (all except geographical diversity) lead to a decrease in redundancy in reviews. Furthermore, publication network-based diversity alone also helps bring in varying perspectives (that is, low redundancy), even within specific review criteria. Our study adopts a group decision-making perspective for reviewer assignments in peer review and suggests dimensions of diversity that can help guide the reviewer assignment process.

Figures

Figures reproduced from arXiv: 2411.11437 by the authors.

Figure 1
Figure 1. Causal Diagram consisting of reviewer-specific features (oval nodes)—reviewer profile and reviewer expertise with [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 32 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Rebecca Abma-Schouten, Joey Gijbels, Wendy Reijmerink, and Ingeborg Meijer. 2023. https://doi.org/10.1093/scipol/scad009 Evaluation of research proposals by peer review panels: broader panels for broader assessments? Science and Public Policy, 50(4):619--632

  4. [4]

    Karen A Bantel and Susan E Jackson. 1989. https://onlinelibrary.wiley.com/doi/abs/10.1002/smj.4250100709 Top management and innovations in banking: Does the composition of the top team make a difference? Strategic management journal, 10(S1):107--124

  5. [5]

    Yoav Benjamini and Yosef Hochberg. 1995. http://www.jstor.org/stable/2346101 Controlling the false discovery rate: A practical and powerful approach to multiple testing . Journal of the Royal Statistical Society. Series B (Methodological), 57(1):289--300

  6. [6]

    Federico Bianchi and Flaminio Squazzoni. 2015. https://dl.acm.org/doi/10.5555/2888619.2889159 Is three better than one? S imulating the effect of reviewer selection and behavior on the quality and efficiency of peer review . In Proceedings of the 2015 Winter Simulation Conference, WSC '15, page 4081–4089. IEEE Press

  7. [7]

    David Blei, Andrew Ng, and Michael Jordan. 2003. https://dl.acm.org/doi/abs/10.5555/944919.944937 Latent dirichlet allocation . The Journal of Machine Learning Research, 3(1):993–1022

  8. [8]

    Harvey Brooks. 1978. http://www.jstor.org/stable/20024552 The problem of research priorities . Daedalus, 107(2):171--190

Show all 52 references
  1. [9]

    Laurent Charlin and Richard Zemel. 2013. https://openreview.net/forum?id=caynafZAnBafx The toronto paper matching system: A n automated paper-reviewer assignment system . In Proceedings of the International Conference on Machine Learning Workshop on Peer Reviewing and Publishi...

  2. [10]

    Daryl E Chubin and Edward J Hackett. 1990. Peerless science: Peer review and US science policy. State University of New York Press

  3. [11]

    Ronald Aylmer Fisher. 1936. https://pmc.ncbi.nlm.nih.gov/articles/PMC2458144/ Design of experiments . British Medical Journal, 1(3923):554

  4. [12]

    a \"a n \

    Mikael Fogelholm, Saara Leppinen, Anssi Auvinen, Jani Raitanen, Anu Nuutinen, and Kalervo V \"a \"a n \"a nen. 2012. https://www.sciencedirect.com/science/article/pii/S089543561100148X Panel discussion does not improve reliability of peer review for medical research grant prop...

  5. [13]

    Francisco Grimaldo and Mario Paolucci. 2013. https://www.worldscientific.com/doi/abs/10.1142/S0219525913500045 A simulation of disagreement for control of rational cheating in peer review . Advances in Complex Systems, 16(07):1350004

  6. [14]

    Lowell Hargens and Jerald Herting. 1990. https://doi.org/10.1007/BF02130467 Neglected considerations in the analysis of agreement among journal referees . Scientometrics, 19(1-2):91--106

  7. [15]

    L Richard Hoffman, Ernest Harburg, and Norman RF Maier. 1962. https://doi.org/10.1037/h0045952 Differences and disagreement as factors in creative group problem solving. The Journal of Abnormal and Social Psychology, 64(3):206

  8. [16]

    Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A

    Bas Hofstra, Vivek V. Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, and Daniel A. McFarland. 2020. https://doi.org/10.1073/pnas.1915378117 The diversity–innovation paradox in science . Proceedings of the National Academy of Sciences, 117(17):9284--9291

  9. [17]

    Lu Hong and Scott E. Page. 2004. https://doi.org/10.1073/pnas.0403723101 Groups of diverse problem solvers can outperform groups of high-ability problem solvers . Proceedings of the National Academy of Sciences, 101(46):16385--16389

  10. [18]

    https://aclanthology.org/N19-1219

    Xinyu Hua, Mitko Nikolov, Nikhil Badugu, and Lu Wang. 2019. "https://aclanthology.org/N19-1219" Argument mining for understanding peer reviews . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language...

  11. [19]

    Susan E Jackson, Karen E May, Kristina Whitney, Richard A Guzzo, and Eduardo Salas. 1995. Understanding the dynamics of diversity in decision-making teams. Team effectiveness and decision making in organizations, 204:261

  12. [20]

    Steven Jecmen, Hanrui Zhang, Ryan Liu, Nihar Shah, Vincent Conitzer, and Fei Fang. 2020. https://proceedings.neurips.cc/paper_files/paper/2020/file/93fb39474c51b8a82a68413e2a5ae17a-Paper.pdf Mitigating manipulation in peer review via randomized reviewer assignments . In Advanc...

  13. [21]

    Tom Jefferson, Philip Alderson, Elizabeth Wager, and Frank Davidoff. 2002 a . https://doi.org/10.1001/jama.287.21.2784 Effects of editorial peer review: a systematic review . JAMA, 287(21):2784--2786

  14. [22]

    Tom Jefferson, Elizabeth Wager, and Frank Davidoff. 2002 b . https://doi.org/10.1001/jama.287.21.2786 Measuring the Quality of Editorial Peer Review . JAMA, 287(21):2786--2790

  15. [23]

    Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan K Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, et al. 2024. http://arxiv.org/abs/2405.06563 What can natural language processing do for peer review? arXiv prepr...

  16. [24]

    Grit Laudel. 2006. https://doi.org/10.3152/147154406781776048 Conclave in the Tower of Babel: H ow peers review interdisciplinary research proposals . Research Evaluation, 15(1):57--68

  17. [25]

    Carole J. Lee. 2012. https://doi.org/10.1086/667841 A kuhnian critique of psychometric research on peer review . Philosophy of Science, 79(5):859–870

  18. [26]

    Carole J Lee, Cassidy R Sugimoto, Guo Zhang, and Blaise Cronin. 2013. https://doi.org/10.1002/asi.22784 Bias in peer review . Journal of the American Society for Information Science and Technology, 64(1):2--17

  19. [27]

    Sheen S Levine, Evan P Apfelbaum, Mark Bernard, Valerie L Bartelt, Edward J Zajac, and David Stark. 2014. https://doi.org/10.1073/pnas.1407301111 Ethnic diversity deflates price bubbles . Proceedings of the National Academy of Sciences, 111(52):18524--18529

  20. [28]

    Kevin Leyton-Brown, Mausam, Yatin Nandwani, Hedayat Zarkoob, Chris Cameron, Neil Newman, and Dinesh Raghu. 2024. https://doi.org/https://doi.org/10.1016/j.artint.2024.104119 Matching papers and reviewers at large conferences . Artificial Intelligence, 331:104119

  21. [29]

    Collings, Jennifer Raymond, and Cassidy R

    Dakota Murray, Kyle Siler, Vincent Larivi \`e re, Wei Mun Chan, Andrew M. Collings, Jennifer Raymond, and Cassidy R. Sugimoto. 2018. https://doi.org/10.1101/400515 Gender and international diversity improves equity in peer review . bioRxiv

  22. [30]

    Michael Obrecht, Karl Tibelius, and Guy D'Aloisio. 2007. https://doi.org/10.3152/095820207X223785 Examining the value added by committee discussion in the review of applications for research awards . Research Evaluation, 16(2):79--91

  23. [31]

    Meike Olbrecht and Lutz Bornmann. 2010. https://doi.org/10.3152/095820210X12809191250762 Panel peer review of grant applications: what do we know from research in social psychology on judgment and decision-making in groups? Research Evaluation, 19(4):293--304

  24. [32]

    Scott Page. 2008. The difference: How the power of diversity creates better groups, firms, schools, and societies-new edition. Princeton University Press

  25. [33]

    Karl Pearson . 1895. https://www.jstor.org/stable/115794 Note on Regression and Inheritance in the Case of Two Parents . Proceedings of the Royal Society of London Series I, 58:240--242

  26. [34]

    Elizabeth Pier, Joshua Raclaw, Anna Kaatz, Markus Brauer, Molly Carnes, Mitchell Nathan, and Cecilia Ford. 2017. https://doi.org/10.1093/reseval/rvw025 Your comments are meaner than your score: s core calibration talk influences intra-and inter-panel variability during scienti...

  27. [35]

    Porter and Frederick A

    Alan L. Porter and Frederick A. Rossini. 1985. https://doi.org/10.1177/016224398501000304 Peer review of interdisciplinary research proposals . Science, Technology, & Human Values, 10(3):33--38

  28. [36]

    Charvi Rastogi, Xiangchen Song, Zhijing Jin, Ivan Stelmakh, Hal Daum \'e III, Kun Zhang, and Nihar B Shah. 2024. https://arxiv.org/abs/2403.01015 A randomized controlled trial on anonymizing reviewers to each other in peer review discussions . arXiv preprint arXiv:2403.01015

  29. [37]

    Nils Reimers and Iryna Gurevych. 2019. https://aclanthology.org/D19-1410/ Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference o...

  30. [38]

    Alison Reynolds and David Lewis. 2017. https://hbr.org/2017/03/teams-solve-problems-faster-when-theyre-more-cognitively-diverse Teams solve problems faster when they’re more cognitively diverse . Harvard Business Review, 30:1--8

  31. [39]

    Michael R\" o der, Andreas Both, and Alexander Hinneburg. 2015. https://doi.org/10.1145/2684822.2685324 Exploring the space of topic coherence measures . In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, WSDM '15, page 399–408, New York, ...

  32. [40]

    Paul R Rosenbaum and Donald B Rubin. 1983. https://doi.org/10.1093/biomet/70.1.41 The central role of the propensity score in observational studies for causal effects . Biometrika, 70(1):41--55

  33. [41]

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. https://arxiv.org/abs/1910.01108 Distil BERT , a distilled version of BERT : smaller, faster, cheaper and lighter . arXiv preprint arXiv:1910.01108

  34. [42]

    Martin Saveski, Steven Jecmen, Nihar Shah, and Johan Ugander. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/b7d795e655c1463d7299688d489e8ef4-Paper-Conference.pdf Counterfactual evaluation of peer-review assignment policies . In Advances in Neural Information...

  35. [43]

    Nihar B. Shah. 2022. https://doi.org/10.1145/3528086 Challenges, experiments, and computational solutions in peer review . Commun. ACM, 65(6):76–87

  36. [44]

    Samuel R Sommers. 2006. https://doi.org/10.1037/0022-3514.90.4.597 On racial diversity and group decision making: identifying multiple effects of racial composition on jury deliberations. Journal of personality and social psychology, 90(4):597

  37. [45]

    Ivan Stelmakh, Charvi Rastogi, Ryan Liu, Shuchi Chawla, Federico Echenique, and Nihar B Shah. 2023 a . https://doi.org/10.1371/journal.pone.0283980 Cite-seeing and reviewing: A study on citation bias in peer review . PLoS ONE, 18(7):e0283980

  38. [46]

    Ivan Stelmakh, Charvi Rastogi, Nihar B Shah, Aarti Singh, and Hal Daum \'e III. 2023 b . https://doi.org/10.1371/journal.pone.0287443 A large scale randomized controlled trial on herding in peer-review discussions . PLoS ONE, 18(7):e0287443

  39. [47]

    Misha Teplitskiy, Hardeep Ranu, Gary S Gray, Michael Menietti, Eva Guinan, and Karim R Lakhani. 2019. https://www.hbs.edu/ris/Publication Do experts listen to other experts? field experimental evidence from scientific peer review

  40. [48]

    Andrew Tomkins, Min Zhang, and William D. Heavlin. 2017. https://doi.org/10.1073/pnas.1707323114 Reviewer bias in single- versus double-blind peer review . Proceedings of the National Academy of Sciences, 114(48):12708--12713

  41. [49]

    Mark Ware. 2008. Peer review: benefits, perceptions and alternatives. Citeseer

  42. [50]

    Wenting Xiong and Diane Litman. 2011. https://aclanthology.org/P11-2088 Automatically predicting peer-review helpfulness . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 502--507, Portland, Oregon,...

  43. [51]

    Weizhe Yuan, Pengfei Liu, and Graham Neubig. 2022. https://doi.org/10.1613/jair.1.12862 Can we automate scientific reviewing? Journal of Artificial Intelligence Research, 75:171--212

  44. [52]

    Zumel Dumlao and Misha Teplitskiy

    James M. Zumel Dumlao and Misha Teplitskiy. 2023. https://doi.org/10.31235/osf.io/754e3 The effect of reviewer geographical diversity on evaluations is reduced by anonymizing submissions . SocArXiv

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.