Pith. sign in

REVIEW 3 major objections 5 minor 54 references

The Viability of Crowdsourcing for RAG Evaluation

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Human pairwise judgments, not LLM judges or overlap metrics, are the reliable gold standard for RAG response evaluation.

desk verdict Valuable crowdsourced RAG evaluation corpus and solid negative findings, but the reliability claim rests on a circular post-hoc exclusion that needs to be unwound. read the letter →

arxiv 2504.15689 v1 pith:UACF3LXN submitted 2025-04-22 cs.IR

classification cs.IR
keywords retrieval-augmentedgenerationcrowdsourcingRAGevaluationpairwisejudgmentsLLM-as-judgereference-basedmetricsutilitydimensionshumanpreference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether ordinary crowd workers can supply the two things RAG evaluation needs — reference responses to compare against, and judgments of response quality — and whether those crowdsourced products are trustworthy. Its central claim is that human pairwise judgments are the reliable and cost-effective source of ground truth: after filtering out low-competency workers and near-identical response pairs, the remaining five-judge majority votes reach agreement levels on par with earlier information-retrieval studies, at a cost of roughly six cents per gold judgment. The same experiments show that LLM-generated responses are judged better than human-written ones on most utility dimensions, that bullet-style responses are preferred over essay and news styles, and that neither LLM-based pairwise judging nor reference-overlap metrics such as BLEU, ROUGE-L, and BERTScore reproduce the human rankings. The authors caution that the single topic set, the single LLM configuration, and the crowd-worker populations limit how far the results generalize.

What carries the argument

The load-bearing machinery is the pairwise judgment protocol combined with a two-stage cleaning step. Each comparison asks five independent workers to choose between two responses on a seven-dimension rubric, with a neutral option except for overall quality; this replicates the finding of prior work that pairwise formats yield more reliable text-quality judgments than pointwise scales. A competency-weighted majority-vote model estimates per-worker reliability, and judgments from the bottom quarter of workers are dropped; pairs that lack a majority vote on most dimensions are marked as minimally differentiable and excluded from the gold set. Per-topic response rankings are then derived with a probabilistic pairwise-comparison model, and these rankings serve as the ground truth against which LLM judgments and reference-overlap metrics are tested.

What would settle it

Take the full 47,320 pairwise judgments, recompute per-topic rankings with no competency filtering and no exclusion of minimally differentiable pairs, and compare the resulting system and style orderings to the paper's cleaned rankings; if the human-versus-LLM advantage or the bullet-style advantage disappears under a plain majority vote, then those findings depend on the cleaning step rather than on the raw crowd signal. A complementary check is to have expert judges label the excluded near-tie pairs and test whether the cleaned gold labels agree with experts on exactly those items.

Watch

Extended reading notes

Core claim

The paper's discovery is that pairwise comparison — showing a worker two responses and asking which is better on a specific dimension, with a neutral option — is the design that makes crowdsourced RAG evaluation work. Across 65 topics, six responses per topic (three human, three LLM) were paired into 1,352 comparison items and judged by five workers each on seven dimensions: topical correctness, logical and stylistic coherence, broad and deep coverage, internal consistency, and overall quality. Agreement among raw crowd judgments starts low, averaging around 0.19, but two targeted corrections — removing the lowest-competency quarter of workers using a competency-weighted voting model, and setting aside about 20% of pairs where five judges could not form a majority — raise agreement to roughly 0.48, comparable to established annotation studies. The paper uses these cleaned labels to rank responses per topic with a probabilistic pairwise-comparison model, and reports that LLM responses are significantly preferred over human-written ones on most dimensions while bullet-style responses beat essay and news styles. Against these human-derived rankings, both the LLM-as-judge judgments and reference-based similarity metrics score poorly, which the paper reads as evidence that judgment-based evaluation with crowdsourced pairwise data is the viable path for RAG.

Load-bearing premise

The reliability conclusion rests on the cleaning step that discards the lowest-competency quarter of workers and about a fifth of response pairs on which five judges could not form a majority; if either exclusion is biased toward or against a response type, the gold labels and every comparison built on them inherit that bias.

Editorial extensions

If this is right

  • RAG benchmark builders can treat crowdsourced pairwise judgments as a practical source of ground truth, at a cost of roughly six cents per gold judgment once redundancy and filtering are accounted for.
  • Reference-based evaluation should not be used to rank RAG systems: across utility dimensions, even the strongest tested overlap metric reached only a moderate correlation with the human-derived rankings, and most values were much lower.
  • Zero-shot LLM judging, at least in the tested configuration, is not a valid substitute for human pairwise judgments, since agreement with gold labels remains poor even when the LLM is highly self-consistent.
  • Pairwise designs are worth their extra cost over pointwise ones: after the same competency correction, pointwise agreement remained around 0.21, less than half the pairwise level.
  • Response style is a real factor in perceived quality, so RAG systems that produce bullet-style output currently hold an advantage over essay- and news-style output on these topics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step the paper does not take is to test whether a few-shot LLM judge fine-tuned or prompted with a sample of these human gold labels approaches human-level agreement; the corpus's public release makes that test directly runnable.
  • Because the gold set excludes the roughly 20% of pairs judges found minimally differentiable, the reported human-versus-LLM quality gap may describe only clearly distinguishable response pairs; in practice, systems whose outputs are all similar in quality might look closer to ties than the headline numbers suggest.
  • The finding that LLMs write less readable text while humans mirror source readability suggests a cheap, testable intervention: instructing generators to simplify syntax and copy more source-like phrasing could close part of the judged quality gap.
  • If the reliability result transfers to other topic sets and languages, crowdsourced pairwise utility judgments could become the calibration data for automated RAG metrics, shifting the debate from whether LLMs can replace humans to how much human-labeled data is needed to tune LLM judges.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces CrowdRAG-25, a corpus of 903 human-written and 903 LLM-generated RAG responses for 301 TREC RAG'24 topics across three discourse styles, together with 47,320 human pairwise judgments and 10,556 LLM pairwise judgments for a subset of 65 topics over seven utility dimensions. It reports analyses of human versus LLM writing behavior, an assessment of crowd judgment reliability using Krippendorff's alpha and MACE-based competency correction, and comparisons of crowd judgments with LLM-as-judge and reference-based evaluation metrics. The headline claim is that human pairwise judgments are reliable and cost-effective for RAG evaluation, whereas LLM-based pairwise judgments and automated reference-based metrics fail to reproduce human preferences. All data and code are released openly.

Significance. The released corpus is a substantial and potentially reusable resource, and the study is unusually transparent in its cost accounting, worker recruitment, spam controls, and interaction logging. The writing-behavior analyses (Section 4) are informative and largely independent of the reliability claim. However, the central reliability claim is currently conditional: it is established only after post hoc exclusions, and the exclusion rule is entangled with the gold-label definition. If the authors add robustness and sensitivity analyses demonstrating that the exclusions do not bias the system rankings, the paper would make a solid contribution to RAG evaluation methodology and to the debate on LLM-as-judge.

major comments (3)
  1. [Section 5.1, Table 4] The reliability claim rests on a circular exclusion. The 'minimally differentiable' split removes pairs that lack a 3-of-5 majority vote on at least 4 of 7 dimensions, and the same majority-vote procedure then defines the gold label for the remaining pairs. Reporting alpha after this exclusion (0.41 to 0.48) as evidence that the gold labels are reliable does not establish reliability on the full judgment pool, and all downstream comparisons in Tables 7-9 inherit this selection. The paper should report balance statistics for the 290 excluded pairs (by response origin, discourse style, topic, and utility dimension), repeat the Bradley-Terry grades, reference-metric correlations, and LLM-judge agreements without the exclusion or under alternative exclusion rules, and show that the rankings are stable. The 30-item expert validation in Section 3.3.2 is too small and deliberately hard to establish neutrality of the exclusion.
  2. [Section 5.1, worker exclusion] The paper states that approximately the lower quarter of workers by MACE competency are removed to obtain the competency-corrected labels, but it does not specify the exact competency threshold, the number of workers removed, or the number of judgments discarded. Without these details, the cost-per-gold-judgment figure in Table 1 and the reliability improvement in Table 4 cannot be reproduced. Please report the full filtering pipeline and a sensitivity analysis over the exclusion percentile, since the headline claim that crowdsourcing is 'cost-effective' depends directly on this threshold.
  3. [Sections 3.3.2 and 5.1, Table 4] The expert validation is not quantitative enough to support the claim that the exclusions are neutral or that there is no systematic task failure. The text says experts demonstrate higher absolute agreement, yet the expert alpha values in Table 4 for the minimally differentiable split are negative (e.g., -0.11 for topical correctness), which appears inconsistent and needs clarification. Moreover, the experts judged only 30 deliberately hard pairs; what is needed is expert-crowd agreement on a random sample of both excluded and retained pairs, with reported agreement coefficients and confidence intervals, before the cleaned gold labels can be treated as ground truth for Tables 7-9.
minor comments (5)
  1. [Section 1, Introduction] The sentence 'no systematic investigation of the capabilities and limitations of human-sourced ground-truth data for retrieval-augmented generation has not done so far' contains a double negative and should be rewritten.
  2. [Table 3] The table rows for the Jaccard coefficient and Spearman's rho are not explicitly labeled beyond the legend symbols below the table; adding row labels such as 'Human vs. LLM', 'Document ranking vs. citation set', and 'Human-cited vs. LLM-cited' would improve readability.
  3. [Table 4] The pointwise column header 'Pntw.' reports nItems = 410, but Section 3.3.1 states that 1,645 pointwise judgments were collected for 47 responses; the relationship between these numbers should be explained.
  4. [Table 9] The column headers in Table 9 are garbled in the rendering, with alpha and rho symbols and subscripts not clearly separated; please reformat the header rows so that each agreement and correlation coefficient is unambiguously labeled.
  5. [Section 5.3, Table 8] The Spearman correlations in Table 8 are reported without significance tests or confidence intervals; given the small per-topic sample size (six responses per topic), a bootstrap or permutation interval would help assess whether the differences across metrics are meaningful.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the paper's reliability and comparison claims are empirical, are transparently conditioned on reported data-cleaning, and do not reduce to their inputs by construction.

full rationale

This is an empirical benchmark paper, not a derivational chain, and I find no step in which a prediction or claimed first-principles result is equivalent to an input by definition. The utility dimensions are imported from the authors' own prior framework [13], and the Bradley-Terry inference from [14], but these are openly cited methodological imports rather than unverified uniqueness theorems or ansatze smuggled in to force the central result. The main reliability claim is explicitly conditional on cleaning: the paper states that "Crowdsourcing can be a reliable source of judgment data for RAG evaluation when controlling for worker competence," and it discloses the raw agreement (mean Krippendorff alpha 0.19 in Table 4) before reporting the post-filtering value of 0.48. The removal of the bottom quarter of workers by MACE competency and of about 20% of pairs classified as "minimally differentiable" is a transparent quality-control step, not a hidden construction: the paper reports the raw agreement, defines the filtering criterion in terms of voting behavior, and separately validates a 30-item sample with experts. The LLM-as-judge comparisons and reference-metric correlations are then evaluated against the resulting human gold labels, but those comparisons are not fitted to or derived from the gold labels; they are external agreement measurements. No equation in the paper reduces to a fitted constant, no fitted parameter is renamed as a prediction, and the self-citations are not load-bearing as unverified evidence for the central empirical claims.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

This is an empirical corpus study; no equations are derived from axioms. The load-bearing assumptions are about data representativeness and the validity of crowd-modeled gold labels. The reliability estimate is sensitive to the hand-chosen exclusion thresholds listed above.

free parameters (3)
  • MACE worker exclusion threshold = bottom ~25% of workers by MACE competency
    Hand-chosen cut used to filter judges; directly raises agreement from 0.19 to 0.41 (Section 5.1, Table 4).
  • Minimally differentiable pair exclusion rule = no majority vote (3 of 5) on at least 4 of 7 dimensions, approximately 20% of pairs
    Post hoc rule to remove pairs with low agreement before computing gold labels and downstream analyses (Section 5.1).
  • Minimum questionnaire time for spam rejection = 10 minutes per questionnaire, ~20 seconds per item
    Hand-set screening threshold that affects which judgments enter the dataset (Section 3.2.3).
assumptions (4)
  • domain assumption TREC RAG'24 topics and webis-01 top-20 passages adequately represent RAG evaluation scenarios.
    Generalization of all findings to other RAG settings relies on this; the authors acknowledge limited generalizability in Section 6.
  • domain assumption The seven utility dimensions from Gienapp et al. [13] capture the relevant aspects of RAG response quality.
    Dimensions are adopted from the authors' own prior framework; their validity is assumed, not independently tested.
  • domain assumption Competency-corrected majority votes (MACE) are valid ground truth for response utility.
    Gold labels are produced by the model, and the paper validates on only 30 expert-judged pairs (Section 3.3.2).
  • domain assumption Pairwise preference grades derived via Bradley-Terry reflect true system ordering.
    Used to rank responses per topic (Section 5.2); relies on standard model assumptions about preferences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Viability of Crowdsourcing for RAG Evaluation." pith.science (2026). https://pith.science/paper/UACF3LXN

@misc{pith2026250415689,
  author       = {Pith},
  title        = {Pith review of: The Viability of Crowdsourcing for RAG Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UACF3LXN}},
  note         = {Machine review of arXiv:2504.15689}
}
read the original abstract

How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. We present the Crowd RAG Corpus 2025 (CrowdRAG-25), which consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC RAG'24 track, across the three discourse styles 'bulleted list', 'essay', and 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. Human pairwise judgments provide reliable and cost-effective results compared to LLM-based pairwise or human/LLM-based pointwise judgments, as well as automated comparisons with human-written reference responses. All our data and tools are freely available.

Figures

Figures reproduced from arXiv: 2504.15689 by the authors.

Figure 1
Figure 1. Distribution of space-separated words, citation [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Left: probability of a document being cited by rank [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Cumulative proportion of statement/reference pairs [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 17 canonical work pages

  1. [13]

    Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang, Johannes Kiesel, Shahbaz Syed, Maik Fröbe, Guido Zuccon, Benno Stein, Matthias Hagen, and Martin Potthast. 2024. Evaluating Generative Ad Hoc Information Retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval...

  2. [1]

    Omar Alonso. 2015. Practical Lessons for Gathering Quality Labels at Scale. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, Santiago, Chile, August 9-13, 2015 , Ricardo Baeza-Yates, Mounia Lalmas, Alistair Moffat, and Berthier A. Ribeiro-Neto (Eds.). ACM, 1089–1092. https://doi.org/10.114...

  3. [2]

    Omar Alonso and Matthew Lease. 2011. Crowdsourcing for information retrieval: principles, methods, and applications. In Proceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011 , Wei-Ying Ma, Jian-Yun Nie, Ricardo Baeza-Yates, Tat-Seng Chua, and W. Bruce Cr...

  4. [3]

    Omar Alonso and Stefano Mizzaro. 2012. Using crowdsourcing for TREC relevance assessment. Inf. Process. Manag. 48, 6 (2012), 1053–1066. https: //doi.org/10.1016/J.IPM.2012.01.004

  5. [4]

    Omar Alonso, Stefano Mizzaro, et al. 2009. Can we get rid of TREC assessors? Using Mechanical Turk for relevance assessment. In Proceedings of the SIGIR 2009 Workshop on the Future of IR Evaluation , Vol. 15. 16

  6. [5]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Inno- vative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intellig...

  7. [6]

    Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All That’s ’Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJC...

  8. [7]

    Charles L. A. Clarke and Laura Dietz. 2024. LLM-based relevance assessment still can’t replace human relevance assessment. CoRR abs/2412.17156 (2024). https://doi.org/10.48550/ARXIV.2412.17156 arXiv:2412.17156

Show all 54 references
  1. [9]

    Shahul ES, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RA- GAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computa- tional Linguistics, EACL 2024 - System Demon...

  2. [10]

    Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023. Perspectives on Large Language Models for Relevance Judgment. In Proceedings o...

  3. [11]

    Robert Friel, Masha Belyi, and Atindriyo Sanyal. 2024. RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems. CoRR abs/2407.11005 (2024). https://doi.org/10.48550/ARXIV.2407.11005 arXiv:2407.11005

  4. [12]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. CoRR abs/2312.10997 (2023). https://doi.org/10.48550/ARXIV.2312.10997 arX...

  5. [14]

    Lukas Gienapp, Benno Stein, Matthias Hagen, and Martin Potthast. 2020. Efficient Pairwise Annotation of Argument Quality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Dan Jurafsky, Joyce Chai, Na...

  6. [15]

    Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News Summarization and Evaluation in the Era of GPT-3. CoRR abs/2209.12356 (2022). https://doi.org/10. 48550/ARXIV.2209.12356 arXiv:2209.12356

  7. [16]

    Matthias Hagen, Martin Potthast, Michael Völske, Jakob Gomoll, and Benno Stein. 2016. How Writers Search: Analyzing the Search and Writing Logs of Non- fictional Essays. InProceedings of the 2016 ACM Conference on Human Information Interaction and Retrieval, CHIIR 2016, Carrbo...

  8. [17]

    Tom Hosking, Phil Blunsom, and Max Bartolo. 2024. Human Feedback is not Gold Standard. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net. https://openreview. net/forum?id=7W3GLNImfS

  9. [18]

    Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard H. Hovy

  10. [19]

    Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023. HAGRID: A Human-LLM Collaborative Dataset for Generative Information- Seeking with Attribution. CoRR abs/2307.16883 (2023). https://doi.org/10.48550/ ARXIV.2307.16883 arXiv:2307.16883

  11. [20]

    J Peter Kincaid, RP Fishburne, RL Rogers, and BS Chissom. 1975. Derivation of new readability formulas (Automated Reliability Index, Fog Count and Flesch Reading Ease Formula) for Navy enlisted personnel (Research Branch Report 8-75). Memphis, TN: Naval Air Station; 1975. Nava...

  12. [21]

    Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...

  13. [22]

    Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2024. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. CoRR abs/2411.16...

  14. [23]

    Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. 2024. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods. CoRR abs/2412.05579 (2024). https://doi.org/10.48550/ ARXIV.2412.05579 arXiv:2412.05579

  15. [24]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013/

  16. [25]

    Jimmy Lin and Dina Demner-Fushman. 2006. Methods for automatically eval- uating answers to complex questions. Inf. Retr. 9, 5 (2006), 565–587. https: //doi.org/10.1007/S10791-006-9003-7

  17. [26]

    Yi Liu, Lianzhe Huang, Shicheng Li, Sishuo Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023. RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge. CoRR abs/2311.08147 (2023). https: //doi.org/10.48550/ARXIV.2311.08147 arXiv:2311.08147

  18. [27]

    Yuanjie Lyu, Zhiyu Li, Simin Niu, Feiyu Xiong, Bo Tang, Wenjin Wang, Hao Wu, Huanyong Liu, Tong Xu, and Enhong Chen. 2024. CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models. CoRR abs/2401.17043 (2024). https://doi.org/10.48...

  19. [28]

    Sean MacAvaney and Luca Soldaini. 2023. One-Shot Labeling for Automatic Rel- evance Estimation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Tai- wan, July 23-27, 2023 , Hsin-Hsi Chen, W...

  20. [29]

    Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. 2018. RankME: Reliable Human Ratings for Natural Language Generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-...

  21. [30]

    Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. InProceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Gr...

  22. [31]

    Martin Potthast, Matthias Hagen, and Benno Stein. 2020. The dilemma of the direct answer. SIGIR Forum 54, 1 (2020), 14:1–14:12. https://doi.org/10.1145/ 3451964.3451978

  23. [32]

    Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track. CoRR abs/2406.16828 (2024). https://doi.org/10...

  24. [33]

    Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework. CoRR abs/2411.09607 (2024). https://doi.org/10.48550/ARXIV.2411.09607 arXiv:2411.09607

  25. [34]

    Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Paul Thomas

    Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Paul Thomas. 2024. SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval. CoRR abs/2408.16312 (2024). https://doi.org/10.48550/ARXIV.2408. 16312 arXiv:2408.16312

  26. [35]

    Rahmani, Emine Yilmaz, Nick Craswell, and Bhaskar Mitra

    Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, and Bhaskar Mitra

  27. [36]

    Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, and Falk Scholer. 2021. On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation. Inf. Process. Manag. 58, 6 (2021), 102688. https://doi.org/10. 1016/J.IPM.2021.102688

  28. [37]

    Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024. ARES: An Automated Evaluation Framework for Retrieval-Augmented Genera- tion Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics...

  29. [38]

    Harrisen Scells, Jimmy, and Guido Zuccon. 2021. Big Brother: A Drop-In Website Interaction Logging Service. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, Fernando Diaz, C...

  30. [39]

    Ian Soboroff. 2024. Don’t Use LLMs to Make Relevance Judgments. CoRR abs/2409.15133 (2024). https://doi.org/10.48550/ARXIV.2409.15133 arXiv:2409.15133

  31. [40]

    Yixuan Tang and Yi Yang. 2024. MultiHop-RAG: Benchmarking Retrieval- Augmented Generation for Multi-Hop Queries. CoRR abs/2401.15391 (2024). https://doi.org/10.48550/ARXIV.2401.15391 arXiv:2401.15391

  32. [41]

    Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024. Large Language Models can Accurately Predict Searcher Preferences. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC,...

  33. [42]

    Mehmet Deniz Türkmen, Mucahid Kutlu, Bahadir Altun, and Gökalp Cosgun

  34. [43]

    Shivani Upadhyay, Ehsan Kamalloo, and Jimmy Lin. 2024. LLMs Can Patch Up Missing Relevance Judgments in Evaluation. CoRR abs/2405.04727 (2024). https://doi.org/10.48550/ARXIV.2405.04727 arXiv:2405.04727

  35. [44]

    Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, and Jimmy Lin. 2024. A Large- Scale Study of Relevance Assessments with Large Language Models: An Initial Look. CoRR abs/2411.08275 (2024). https://doi.org/10.48550/ARXI...

  36. [45]

    Bernstein

    Vasilis Verroios and Michael S. Bernstein. 2014. Context Trees: Crowdsourcing Global Understanding from Local Views. In Proceedings of the Seconf AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2014, November 2-4, 2014, Pittsburgh, Pennsylvania, USA , Jeffrey P. ...

  37. [46]

    Voorhees

    Ellen M. Voorhees. 2003. Overview of the TREC 2003 Question Answering Track. In Proceedings of The Twelfth Text REtrieval Conference, TREC 2003, Gaithersburg, Maryland, USA, November 18-21, 2003 (NIST Special Publication, Vol. 500-255) , Ellen M. Voorhees and Lori P. Buckland ...

  38. [47]

    Shuting Wang, Jiongnan Liu, Shiren Song, Jiehan Cheng, Yuqi Fu, Peidong Guo, Kun Fang, Yutao Zhu, and Zhicheng Dou. 2024. DomainRAG: A Chi- nese Benchmark for Evaluating Domain-specific Retrieval-Augmented Genera- tion. CoRR abs/2406.05654 (2024). https://doi.org/10.48550/ARXI...

  39. [48]

    Meng-Han Wu and Alexander J. Quinn. 2017. Confusing the Crowd: Task Instruction Quality on Amazon Mechanical Turk. In Proceedings of the Fifth AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2017, 23-26 October 2017, Québec City, Québec, Canada , Steven Dow and A...

  40. [49]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi

  41. [50]

    McKeown, and Tatsunori B

    Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen R. McKeown, and Tatsunori B. Hashimoto. 2024. Benchmarking Large Language Models for News Summarization. Trans. Assoc. Comput. Linguistics 12 (2024), 39–57. https://doi.org/10.1162/TACL\_A\_00632

  42. [51]

    Weijia Zhang, Mohammad Aliannejadi, Yifei Yuan, Jiahuan Pei, Jia-Hong Huang, and Evangelos Kanoulas. 2024. Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics. In Proceedings of the 17th International Natural Language Gene...

  43. [2013]

    Learning Whom to Trust with MACE. In Human Language Technologies: Conference of the North American Chapter of the Association of Computational Linguistics, Proceedings, June 9-14, 2013, Westin Peachtree Plaza Hotel, Atlanta, Georgia, USA, Lucy Vanderwende, Hal Daumé III, and K...

  44. [2020]

    In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020

    BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. https://openreview.net/forum?id=SkeHuCVFDr

  45. [2024]

    CoRR abs/2412.13268 (2024)

    JudgeBlender: Ensembling Judgments for Automatic Relevance Assess- ment. CoRR abs/2412.13268 (2024). https://doi.org/10.48550/ARXIV.2412.13268 arXiv:2412.13268

  46. [2025]

    CoRR abs/2501.02408 (2025)

    GenTREC: The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems. CoRR abs/2501.02408 (2025). https://doi.org/10.48550/ARXIV.2501.02408 arXiv:2501.02408

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.