REVIEW 3 major objections 5 minor 54 references
The Viability of Crowdsourcing for RAG Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Human pairwise judgments, not LLM judges or overlap metrics, are the reliable gold standard for RAG response evaluation.
desk verdict Valuable crowdsourced RAG evaluation corpus and solid negative findings, but the reliability claim rests on a circular post-hoc exclusion that needs to be unwound. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the pairwise judgment protocol combined with a two-stage cleaning step. Each comparison asks five independent workers to choose between two responses on a seven-dimension rubric, with a neutral option except for overall quality; this replicates the finding of prior work that pairwise formats yield more reliable text-quality judgments than pointwise scales. A competency-weighted majority-vote model estimates per-worker reliability, and judgments from the bottom quarter of workers are dropped; pairs that lack a majority vote on most dimensions are marked as minimally differentiable and excluded from the gold set. Per-topic response rankings are then derived with a probabilistic pairwise-comparison model, and these rankings serve as the ground truth against which LLM judgments and reference-overlap metrics are tested.
What would settle it
Take the full 47,320 pairwise judgments, recompute per-topic rankings with no competency filtering and no exclusion of minimally differentiable pairs, and compare the resulting system and style orderings to the paper's cleaned rankings; if the human-versus-LLM advantage or the bullet-style advantage disappears under a plain majority vote, then those findings depend on the cleaning step rather than on the raw crowd signal. A complementary check is to have expert judges label the excluded near-tie pairs and test whether the cleaned gold labels agree with experts on exactly those items.
Extended reading notes
Core claim
The paper's discovery is that pairwise comparison — showing a worker two responses and asking which is better on a specific dimension, with a neutral option — is the design that makes crowdsourced RAG evaluation work. Across 65 topics, six responses per topic (three human, three LLM) were paired into 1,352 comparison items and judged by five workers each on seven dimensions: topical correctness, logical and stylistic coherence, broad and deep coverage, internal consistency, and overall quality. Agreement among raw crowd judgments starts low, averaging around 0.19, but two targeted corrections — removing the lowest-competency quarter of workers using a competency-weighted voting model, and setting aside about 20% of pairs where five judges could not form a majority — raise agreement to roughly 0.48, comparable to established annotation studies. The paper uses these cleaned labels to rank responses per topic with a probabilistic pairwise-comparison model, and reports that LLM responses are significantly preferred over human-written ones on most dimensions while bullet-style responses beat essay and news styles. Against these human-derived rankings, both the LLM-as-judge judgments and reference-based similarity metrics score poorly, which the paper reads as evidence that judgment-based evaluation with crowdsourced pairwise data is the viable path for RAG.
Load-bearing premise
The reliability conclusion rests on the cleaning step that discards the lowest-competency quarter of workers and about a fifth of response pairs on which five judges could not form a majority; if either exclusion is biased toward or against a response type, the gold labels and every comparison built on them inherit that bias.
Editorial extensions
If this is right
- RAG benchmark builders can treat crowdsourced pairwise judgments as a practical source of ground truth, at a cost of roughly six cents per gold judgment once redundancy and filtering are accounted for.
- Reference-based evaluation should not be used to rank RAG systems: across utility dimensions, even the strongest tested overlap metric reached only a moderate correlation with the human-derived rankings, and most values were much lower.
- Zero-shot LLM judging, at least in the tested configuration, is not a valid substitute for human pairwise judgments, since agreement with gold labels remains poor even when the LLM is highly self-consistent.
- Pairwise designs are worth their extra cost over pointwise ones: after the same competency correction, pointwise agreement remained around 0.21, less than half the pairwise level.
- Response style is a real factor in perceived quality, so RAG systems that produce bullet-style output currently hold an advantage over essay- and news-style output on these topics.
Reading between the lines
- A natural next step the paper does not take is to test whether a few-shot LLM judge fine-tuned or prompted with a sample of these human gold labels approaches human-level agreement; the corpus's public release makes that test directly runnable.
- Because the gold set excludes the roughly 20% of pairs judges found minimally differentiable, the reported human-versus-LLM quality gap may describe only clearly distinguishable response pairs; in practice, systems whose outputs are all similar in quality might look closer to ties than the headline numbers suggest.
- The finding that LLMs write less readable text while humans mirror source readability suggests a cheap, testable intervention: instructing generators to simplify syntax and copy more source-like phrasing could close part of the judged quality gap.
- If the reliability result transfers to other topic sets and languages, crowdsourced pairwise utility judgments could become the calibration data for automated RAG metrics, shifting the debate from whether LLMs can replace humans to how much human-labeled data is needed to tune LLM judges.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CrowdRAG-25, a corpus of 903 human-written and 903 LLM-generated RAG responses for 301 TREC RAG'24 topics across three discourse styles, together with 47,320 human pairwise judgments and 10,556 LLM pairwise judgments for a subset of 65 topics over seven utility dimensions. It reports analyses of human versus LLM writing behavior, an assessment of crowd judgment reliability using Krippendorff's alpha and MACE-based competency correction, and comparisons of crowd judgments with LLM-as-judge and reference-based evaluation metrics. The headline claim is that human pairwise judgments are reliable and cost-effective for RAG evaluation, whereas LLM-based pairwise judgments and automated reference-based metrics fail to reproduce human preferences. All data and code are released openly.
Significance. The released corpus is a substantial and potentially reusable resource, and the study is unusually transparent in its cost accounting, worker recruitment, spam controls, and interaction logging. The writing-behavior analyses (Section 4) are informative and largely independent of the reliability claim. However, the central reliability claim is currently conditional: it is established only after post hoc exclusions, and the exclusion rule is entangled with the gold-label definition. If the authors add robustness and sensitivity analyses demonstrating that the exclusions do not bias the system rankings, the paper would make a solid contribution to RAG evaluation methodology and to the debate on LLM-as-judge.
major comments (3)
- [Section 5.1, Table 4] The reliability claim rests on a circular exclusion. The 'minimally differentiable' split removes pairs that lack a 3-of-5 majority vote on at least 4 of 7 dimensions, and the same majority-vote procedure then defines the gold label for the remaining pairs. Reporting alpha after this exclusion (0.41 to 0.48) as evidence that the gold labels are reliable does not establish reliability on the full judgment pool, and all downstream comparisons in Tables 7-9 inherit this selection. The paper should report balance statistics for the 290 excluded pairs (by response origin, discourse style, topic, and utility dimension), repeat the Bradley-Terry grades, reference-metric correlations, and LLM-judge agreements without the exclusion or under alternative exclusion rules, and show that the rankings are stable. The 30-item expert validation in Section 3.3.2 is too small and deliberately hard to establish neutrality of the exclusion.
- [Section 5.1, worker exclusion] The paper states that approximately the lower quarter of workers by MACE competency are removed to obtain the competency-corrected labels, but it does not specify the exact competency threshold, the number of workers removed, or the number of judgments discarded. Without these details, the cost-per-gold-judgment figure in Table 1 and the reliability improvement in Table 4 cannot be reproduced. Please report the full filtering pipeline and a sensitivity analysis over the exclusion percentile, since the headline claim that crowdsourcing is 'cost-effective' depends directly on this threshold.
- [Sections 3.3.2 and 5.1, Table 4] The expert validation is not quantitative enough to support the claim that the exclusions are neutral or that there is no systematic task failure. The text says experts demonstrate higher absolute agreement, yet the expert alpha values in Table 4 for the minimally differentiable split are negative (e.g., -0.11 for topical correctness), which appears inconsistent and needs clarification. Moreover, the experts judged only 30 deliberately hard pairs; what is needed is expert-crowd agreement on a random sample of both excluded and retained pairs, with reported agreement coefficients and confidence intervals, before the cleaned gold labels can be treated as ground truth for Tables 7-9.
minor comments (5)
- [Section 1, Introduction] The sentence 'no systematic investigation of the capabilities and limitations of human-sourced ground-truth data for retrieval-augmented generation has not done so far' contains a double negative and should be rewritten.
- [Table 3] The table rows for the Jaccard coefficient and Spearman's rho are not explicitly labeled beyond the legend symbols below the table; adding row labels such as 'Human vs. LLM', 'Document ranking vs. citation set', and 'Human-cited vs. LLM-cited' would improve readability.
- [Table 4] The pointwise column header 'Pntw.' reports nItems = 410, but Section 3.3.1 states that 1,645 pointwise judgments were collected for 47 responses; the relationship between these numbers should be explained.
- [Table 9] The column headers in Table 9 are garbled in the rendering, with alpha and rho symbols and subscripts not clearly separated; please reformat the header rows so that each agreement and correlation coefficient is unambiguously labeled.
- [Section 5.3, Table 8] The Spearman correlations in Table 8 are reported without significance tests or confidence intervals; given the small per-topic sample size (six responses per topic), a bootstrap or permutation interval would help assess whether the differences across metrics are meaningful.
Circularity Check
No significant circularity; the paper's reliability and comparison claims are empirical, are transparently conditioned on reported data-cleaning, and do not reduce to their inputs by construction.
full rationale
This is an empirical benchmark paper, not a derivational chain, and I find no step in which a prediction or claimed first-principles result is equivalent to an input by definition. The utility dimensions are imported from the authors' own prior framework [13], and the Bradley-Terry inference from [14], but these are openly cited methodological imports rather than unverified uniqueness theorems or ansatze smuggled in to force the central result. The main reliability claim is explicitly conditional on cleaning: the paper states that "Crowdsourcing can be a reliable source of judgment data for RAG evaluation when controlling for worker competence," and it discloses the raw agreement (mean Krippendorff alpha 0.19 in Table 4) before reporting the post-filtering value of 0.48. The removal of the bottom quarter of workers by MACE competency and of about 20% of pairs classified as "minimally differentiable" is a transparent quality-control step, not a hidden construction: the paper reports the raw agreement, defines the filtering criterion in terms of voting behavior, and separately validates a 30-item sample with experts. The LLM-as-judge comparisons and reference-metric correlations are then evaluated against the resulting human gold labels, but those comparisons are not fitted to or derived from the gold labels; they are external agreement measurements. No equation in the paper reduces to a fitted constant, no fitted parameter is renamed as a prediction, and the self-citations are not load-bearing as unverified evidence for the central empirical claims.
Assumptions & free parameters
free parameters (3)
- MACE worker exclusion threshold =
bottom ~25% of workers by MACE competency
- Minimally differentiable pair exclusion rule =
no majority vote (3 of 5) on at least 4 of 7 dimensions, approximately 20% of pairs
- Minimum questionnaire time for spam rejection =
10 minutes per questionnaire, ~20 seconds per item
assumptions (4)
- domain assumption TREC RAG'24 topics and webis-01 top-20 passages adequately represent RAG evaluation scenarios.
- domain assumption The seven utility dimensions from Gienapp et al. [13] capture the relevant aspects of RAG response quality.
- domain assumption Competency-corrected majority votes (MACE) are valid ground truth for response utility.
- domain assumption Pairwise preference grades derived via Bradley-Terry reflect true system ordering.
Cite this review
Pith. "Pith review of The Viability of Crowdsourcing for RAG Evaluation." pith.science (2026). https://pith.science/paper/UACF3LXN
@misc{pith2026250415689,
author = {Pith},
title = {Pith review of: The Viability of Crowdsourcing for RAG Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/UACF3LXN}},
note = {Machine review of arXiv:2504.15689}
}
read the original abstract
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. We present the Crowd RAG Corpus 2025 (CrowdRAG-25), which consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC RAG'24 track, across the three discourse styles 'bulleted list', 'essay', and 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. Human pairwise judgments provide reliable and cost-effective results compared to LLM-based pairwise or human/LLM-based pointwise judgments, as well as automated comparisons with human-written reference responses. All our data and tools are freely available.
Figures
Reference graph
Works this paper leans on
-
[13]
Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang, Johannes Kiesel, Shahbaz Syed, Maik Fröbe, Guido Zuccon, Benno Stein, Matthias Hagen, and Martin Potthast. 2024. Evaluating Generative Ad Hoc Information Retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval...
arXiv 2024
-
[1]
Omar Alonso. 2015. Practical Lessons for Gathering Quality Labels at Scale. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, Santiago, Chile, August 9-13, 2015 , Ricardo Baeza-Yates, Mounia Lalmas, Alistair Moffat, and Berthier A. Ribeiro-Neto (Eds.). ACM, 1089–1092. https://doi.org/10.114...
-
[2]
Omar Alonso and Matthew Lease. 2011. Crowdsourcing for information retrieval: principles, methods, and applications. In Proceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011 , Wei-Ying Ma, Jian-Yun Nie, Ricardo Baeza-Yates, Tat-Seng Chua, and W. Bruce Cr...
arXiv 2011
-
[3]
Omar Alonso and Stefano Mizzaro. 2012. Using crowdsourcing for TREC relevance assessment. Inf. Process. Manag. 48, 6 (2012), 1053–1066. https: //doi.org/10.1016/J.IPM.2012.01.004
-
[4]
Omar Alonso, Stefano Mizzaro, et al. 2009. Can we get rid of TREC assessors? Using Mechanical Turk for relevance assessment. In Proceedings of the SIGIR 2009 Workshop on the Future of IR Evaluation , Vol. 15. 16
work page 2009
-
[5]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Inno- vative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intellig...
2024
-
[6]
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All That’s ’Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJC...
2021
-
[7]
Charles L. A. Clarke and Laura Dietz. 2024. LLM-based relevance assessment still can’t replace human relevance assessment. CoRR abs/2412.17156 (2024). https://doi.org/10.48550/ARXIV.2412.17156 arXiv:2412.17156
Show all 54 references
-
[9]
Shahul ES, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RA- GAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computa- tional Linguistics, EACL 2024 - System Demon...
2024
-
[10]
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023. Perspectives on Large Language Models for Relevance Judgment. In Proceedings o...
2023
- [11]
- [12]
-
[14]
Lukas Gienapp, Benno Stein, Matthias Hagen, and Martin Potthast. 2020. Efficient Pairwise Annotation of Argument Quality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Dan Jurafsky, Joyce Chai, Na...
2020
- [15]
-
[16]
Matthias Hagen, Martin Potthast, Michael Völske, Jakob Gomoll, and Benno Stein. 2016. How Writers Search: Analyzing the Search and Writing Logs of Non- fictional Essays. InProceedings of the 2016 ACM Conference on Human Information Interaction and Retrieval, CHIIR 2016, Carrbo...
2016
-
[17]
Tom Hosking, Phil Blunsom, and Max Bartolo. 2024. Human Feedback is not Gold Standard. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net. https://openreview. net/forum?id=7W3GLNImfS
2024
-
[18]
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard H. Hovy
- [19]
-
[20]
J Peter Kincaid, RP Fishburne, RL Rogers, and BS Chissom. 1975. Derivation of new readability formulas (Automated Reliability Index, Fog Count and Flesch Reading Ease Formula) for Navy enlisted personnel (Research Branch Report 8-75). Memphis, TN: Naval Air Station; 1975. Nava...
1975
-
[21]
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...
2020
-
[22]
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2024. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. CoRR abs/2411.16...
2024 doi
- [23]
-
[24]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013/
2004
-
[25]
Jimmy Lin and Dina Demner-Fushman. 2006. Methods for automatically eval- uating answers to complex questions. Inf. Retr. 9, 5 (2006), 565–587. https: //doi.org/10.1007/S10791-006-9003-7
2006 doi
- [26]
- [27]
-
[28]
Sean MacAvaney and Luca Soldaini. 2023. One-Shot Labeling for Automatic Rel- evance Estimation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Tai- wan, July 23-27, 2023 , Hsin-Hsi Chen, W...
2023
-
[29]
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. 2018. RankME: Reliable Human Ratings for Natural Language Generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-...
2018 doi
-
[30]
Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. InProceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Gr...
2018
-
[31]
Martin Potthast, Matthias Hagen, and Benno Stein. 2020. The dilemma of the direct answer. SIGIR Forum 54, 1 (2020), 14:1–14:12. https://doi.org/10.1145/ 3451964.3451978
2020
- [32]
- [33]
-
[34]
Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Paul Thomas
Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Paul Thomas. 2024. SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval. CoRR abs/2408.16312 (2024). https://doi.org/10.48550/ARXIV.2408. 16312 arXiv:2408.16312
-
[35]
Rahmani, Emine Yilmaz, Nick Craswell, and Bhaskar Mitra
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, and Bhaskar Mitra
-
[36]
Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, and Falk Scholer. 2021. On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation. Inf. Process. Manag. 58, 6 (2021), 102688. https://doi.org/10. 1016/J.IPM.2021.102688
2021
-
[37]
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024. ARES: An Automated Evaluation Framework for Retrieval-Augmented Genera- tion Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics...
2024
-
[38]
Harrisen Scells, Jimmy, and Guido Zuccon. 2021. Big Brother: A Drop-In Website Interaction Logging Service. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, Fernando Diaz, C...
2021
- [39]
- [40]
-
[41]
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024. Large Language Models can Accurately Predict Searcher Preferences. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC,...
2024
-
[42]
Mehmet Deniz Türkmen, Mucahid Kutlu, Bahadir Altun, and Gökalp Cosgun
- [43]
- [44]
-
[45]
Bernstein
Vasilis Verroios and Michael S. Bernstein. 2014. Context Trees: Crowdsourcing Global Understanding from Local Views. In Proceedings of the Seconf AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2014, November 2-4, 2014, Pittsburgh, Pennsylvania, USA , Jeffrey P. ...
2014 doi
-
[46]
Voorhees
Ellen M. Voorhees. 2003. Overview of the TREC 2003 Question Answering Track. In Proceedings of The Twelfth Text REtrieval Conference, TREC 2003, Gaithersburg, Maryland, USA, November 18-21, 2003 (NIST Special Publication, Vol. 500-255) , Ellen M. Voorhees and Lori P. Buckland ...
2003
- [47]
-
[48]
Meng-Han Wu and Alexander J. Quinn. 2017. Confusing the Crowd: Task Instruction Quality on Amazon Mechanical Turk. In Proceedings of the Fifth AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2017, 23-26 October 2017, Québec City, Québec, Canada , Steven Dow and A...
2017 doi
-
[49]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[50]
McKeown, and Tatsunori B
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen R. McKeown, and Tatsunori B. Hashimoto. 2024. Benchmarking Large Language Models for News Summarization. Trans. Assoc. Comput. Linguistics 12 (2024), 39–57. https://doi.org/10.1162/TACL\_A\_00632
2024 doi
-
[51]
Weijia Zhang, Mohammad Aliannejadi, Yifei Yuan, Jiahuan Pei, Jia-Hong Huang, and Evangelos Kanoulas. 2024. Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics. In Proceedings of the 17th International Natural Language Gene...
2024
-
[2013]
Learning Whom to Trust with MACE. In Human Language Technologies: Conference of the North American Chapter of the Association of Computational Linguistics, Proceedings, June 9-14, 2013, Westin Peachtree Plaza Hotel, Atlanta, Georgia, USA, Lucy Vanderwende, Hal Daumé III, and K...
2013
-
[2020]
In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020
BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. https://openreview.net/forum?id=SkeHuCVFDr
2020
- [2024]
- [2025]
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.