REVIEW 4 major objections 6 minor 1 cited by
EnronQA: Towards Personalized RAG over Private Documents
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read EnronQA, a benchmark of 103,638 private emails and 528,304 question-answer pairs across 150 inboxes, is calibrated so end-to-end RAG accuracy tracks retrieval quality rather than LLM memorization, with no-context accuracy below 5% and…
desk verdict Large, useful private-document RAG benchmark; the calibration claim is partly self-enforcing and human validation is thin, but the resource deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's working parts are the four-gate question admission pipeline and the per-inbox segmentation of the corpus. A question enters the benchmark only if it is specific, objective, grounded, and high quality by automated checks; this is what keeps no-context accuracy near zero and makes end-to-end accuracy move with retrieval recall. The 150 distinct inboxes are what let the same corpus support personalized retrieval experiments rather than only global retrieval.
What would settle it
Have independent human annotators answer a random sample of roughly 500 EnronQA questions with no email in context; if more than a small percentage are answered correctly, the groundedness check is too lenient and the benchmark would not track retrieval quality. The same test can be rerun on each new generation of large language models, since the no-context baseline rising above 5% would signal that the corpus has leaked into training.
Extended reading notes
Core claim
EnronQA is constructed by filtering the cleaned Enron email corpus down to 103,638 usable emails and generating 528,304 QA pairs with a multi-stage LLM pipeline. Each question must pass four automated checks before entering the benchmark: specificity, objectivity, groundedness, and a rules-based quality gate. The central discovery is about calibration: in simulated Recall@1 sweeps, EnronQA is the only benchmark tested where adding the correct context always improves accuracy over the no-context baseline, with no-context accuracy below 5% and nearly a 0.6% accuracy gain per 1% recall gain. The paper also reports baseline RAG results and a memorization case study in which LoRA adapters trained on up to 20,000 facts match long-context performance, while retrieval still outperforms both.
Load-bearing premise
The load-bearing premise is that the four automated checks performed by large language models correctly certify question quality, so the 528,304 question-answer pairs are genuinely answerable from their paired email and not from general knowledge.
Editorial extensions
If this is right
- End-to-end accuracy on EnronQA can be read as a direct measure of retriever quality, since every point of recall adds roughly 0.6 points of accuracy.
- Private-document RAG pipelines can be benchmarked without first controlling for whether the LLM already memorized the documents, which is not possible on Wikipedia-based benchmarks.
- The benchmark's 150-inbox structure enables personalized retrieval evaluation, where systems must find and reason over the right user's documents.
- The memorization case study shows LoRA adapters can recall up to 20,000 facts at a level comparable to putting the facts in context, while retrieval still outperforms both.
- The benchmark is large enough for fine-tuning and continued-pretraining experiments, not just for evaluation.
Reading between the lines
- A consequence the paper leaves implicit: the specificity gate likely pushes questions toward distinctive named entities, which would explain the strong lexical-retrieval results; re-generating questions with paraphrased, entity-light wording would test whether dense retrieval closes the gap.
- The calibration claim is time-sensitive; as future large language models train on broader data, the Enron corpus may leak into training, so re-running the no-context baseline on each new model generation is a simple contamination monitor.
- The memorization results suggest an open capacity-versus-cost tradeoff; because LoRA adapters hold up to 20,000 facts, a natural next experiment is measuring at what fact count a large-rank adapter becomes cheaper or faster than maintaining a search index.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EnronQA, a synthetic question-answering benchmark built from the cleaned Enron email corpus, containing 103,638 emails and 528,304 QA pairs across 150 user inboxes. The dataset is constructed through a DSPy-based pipeline that generates questions with Llama-3.1-70B and filters them through four LLM-based checks: specificity, objectivity, groundedness, and a rule-based quality judge. The authors present dataset statistics, a calibration study comparing EnronQA to NaturalQuestions and TriviaQA, baseline RAG experiments with BM25 and ColBERTv2 over several LLMs, and a case study on LoRA-based memorization of facts. The paper's central claims are that EnronQA is the only evaluated benchmark where adding retrieved context always improves over a no-context baseline, and that it provides a better-calibrated testbed for retriever quality than Wikipedia-based benchmarks.
Significance. If the quality of the 528k synthetic QA pairs is adequately validated, EnronQA would be a valuable community resource: it is large, segmented by user, drawn from realistic private emails, and released with prompts and artifacts. The calibration result, if it holds under independent evaluation, would be useful for the RAG community because it addresses the memorization problem that afflicts Wikipedia-based benchmarks. The LoRA memorization study is also a relevant initial exploration. The paper is transparent about its pipeline and releases substantial supporting material, including prompts and chain-of-thought outputs, which strengthens reproducibility. The main weaknesses are that the evidence for question quality is thin (41 human-labeled-quality examples and 200 judge agreement examples), the evaluation stack is heavily concentrated in one model family, and the headline calibration claim is partly enforced by construction rather than observed independently.
major comments (4)
- [§4.1, Table 3] Table 3 reports an average of 491.81 emails per user, but 103,638 emails divided by 150 users equals 690.92. This arithmetic inconsistency means the summary statistics as reported cannot be trusted until the table is corrected or the discrepancy is explained. Since the size and composition of the benchmark are central to the paper's claims, this needs to be fixed before publication.
- [§4.2 vs. §3.2] The headline calibration claim that 'EnronQA is the only benchmark where adding context is always better than the no-context baseline' is partly enforced by the groundedness filter in §3.2: a question is admitted only when both Llama-3.1-70B and Mixtral-8x7B fail to reproduce the gold answer without the email. The no-context accuracy below 5% is therefore a property of the benchmark's construction, not an independent discovery about private documents. The comparison to NaturalQuestions and TriviaQA, which were not filtered this way, should be framed accordingly, and the authors should report calibration on an unfiltered or differently filtered subset to support the stronger interpretation.
- [§3.2, §4.2, §5.2] The validation of question quality is too thin for a resource of 528k pairs. The LLM judge's 0.98 F1 was measured on only 200 generated match/no-match judgements, and this measures answer-equivalence, not whether a gold answer is factually correct or whether a question is answerable. The quality ruleset was tuned on 21 development questions and tested on 20, with an inter-annotator Spearman correlation of only 0.5. Because the same Llama-3.1-70B family generates questions, filters them, and scores final accuracy, there is a real risk of self-consistency bias. The authors should provide independent human evaluation on a larger random sample, report agreement statistics, and check for unanswerable or ambiguous questions before presenting EnronQA as 'high quality.'
- [§6, Table 5] The claim that LoRA memorization 'matches long-context performance at almost all scales' is not well supported by Table 5. At 5,000 facts the rank-1024 adapter collapses to 0.03 accuracy, and at 20,000 facts ranks 512, 1024, and 2048 yield 0.08, 0.00, and 0.03, respectively. Several cells show non-monotonic behavior across ranks, and no error bars, multiple seeds, or standard deviations are reported. This makes the conclusion about 'surprising capacity' fragile and should be either supported with variance estimates or substantially tempered.
minor comments (6)
- [§1] The text contains an unresolved placeholder '(??)' in the sentence about headroom for improving retriever quality; please replace it with the intended citation or number.
- [Fig. 2 caption] The caption refers to 'Mixtral-7B-Instruct' while the text and reference correctly identify the model as Mixtral-8x7B-Instruct; please correct the model name.
- [§5.2] There is a typo 'GPTo' that should read 'GPT-4o' or 'GPT4o' consistently.
- [Table 4] The column headers in Table 4 are difficult to parse because 'R@5' appears only under 'Query Rewrite' and the row labels mix retriever names with recall values. Please restructure the table so that each retriever's Recall@5 is clearly associated with the right column.
- [Appendix B.7] The introductory sentence in B.7 is duplicated from B.5 and says 'These prompts are used to both answer the question given the context of an email or to produce an answer to the question with no grounding,' which does not describe the quality-evaluation prompt in B.7; please correct it.
- [§6.1] The memorization experiment reports '10 epochs with rate 1e-4' but does not specify the optimizer, learning-rate schedule, or how accuracy is computed from the LLM judge's outputs; please add these details.
Circularity Check
Calibration claim is partly enforced by the groundedness filter: questions answerable without context are removed by construction, so the low no-context baseline is a selection effect rather than an independent discovery.
-
self definitional
[§3.2 QA Generation Pipeline (Groundedness); §4.2 Calibration]
"Groundedness. We determine a question to be "grounded" if neither the llama nor mixtral model can answer the same as the gold answer given no email as context. This both tests that the answers to our questions are not memorized and that the questions are not easily guessable. ... We find that EnronQA is the only benchmark where adding context is always better than the no-context baseline."
The groundedness criterion is an inclusion rule: a QA pair enters EnronQA only when both Llama-3.1-70B and Mixtral-8x7B, given no email, fail to produce the gold answer. The calibration section then measures no-context accuracy on exactly this filtered set and reports it is below 5% and that context always helps. That outcome is logically forced by the filter: any question the same judge could answer without context was discarded. Calling this "the knowledge in EnronQA has not been memorized" restates the admission rule rather than reporting an independent discovery. The Wikipedia comparison is not apples-to-apples, because those benchmarks were not filtered to exclude memorized or guessable questions.
full rationale
EnronQA itself is a corpus-derived resource, not a derivation, and its construction is largely independent: the 103,638-email corpus, filtering steps, and the public Enron source provide external grounding. The self-citation to DSPy/MIPROv2 (Ref. [55], an author of the present paper) is methodological and not load-bearing for the benchmark's validity. The main circularity is narrower: the headline calibration result ('only benchmark where adding context is always better', no-context below 5%) is partly enforced by the groundedness filter in §3.2, which removes questions that the same LLM family can answer without context. The benchmark can still be a useful resource, but the claimed advantage over Wikipedia benchmarks is a designed property, not a discovered one. The QA-quality evidence (0.98 F1 on 200 judge instances, ruleset on 41 labeled questions, same Llama-3.1-70B family for generation and final scoring) is thin but concerns validation strength, not derivation circularity.
Assumptions & free parameters
free parameters (3)
- Corpus filtering thresholds =
Jaccard 0.9, 9 bands, 27 rows; 50-1000 words; length 3-10; ratio >=0.65; ellipsis <10%; confidence 80-90%
- Specificity hard-negative count =
k=10
- LoRA training hyperparameters =
ranks 8-2048; 10 epochs; lr 1e-4; alpha=4x rank; dropout 0.05
assumptions (5)
- domain assumption The 2015 Enron release is a realistic proxy for private enterprise email documents.
- domain assumption The LLM judge correctly determines answer equivalence.
- ad hoc to paper The four quality criteria (specificity, objectivity, groundedness, quality) as implemented cover the relevant notion of question quality.
- domain assumption No-context failure by Llama and Mixtral implies the answer is not in the parametric knowledge of the LLMs tested.
- domain assumption DSPy optimization and synthetic generation produce a diverse question distribution.
Cite this review
Pith. "Pith review of EnronQA: Towards Personalized RAG over Private Documents." pith.science (2026). https://pith.science/paper/MG4SEZAH
@misc{pith2026250500263,
author = {Pith},
title = {Pith review of: EnronQA: Towards Personalized RAG over Private Documents},
year = {2026},
howpublished = {\url{https://pith.science/paper/MG4SEZAH}},
note = {Machine review of arXiv:2505.00263}
}
read the original abstract
Retrieval Augmented Generation (RAG) has become one of the most popular methods for bringing knowledge-intensive context to large language models (LLM) because of its ability to bring local context at inference time without the cost or data leakage risks associated with fine-tuning. A clear separation of private information from the LLM training has made RAG the basis for many enterprise LLM workloads as it allows the company to augment LLM's understanding using customers' private documents. Despite its popularity for private documents in enterprise deployments, current RAG benchmarks for validating and optimizing RAG pipelines draw their corpora from public data such as Wikipedia or generic web pages and offer little to no personal context. Seeking to empower more personal and private RAG we release the EnronQA benchmark, a dataset of 103,638 emails with 528,304 question-answer pairs across 150 different user inboxes. EnronQA enables better benchmarking of RAG pipelines over private data and allows for experimentation on the introduction of personalized retrieval settings over realistic data. Finally, we use EnronQA to explore the tradeoff in memorization and retrieval when reasoning over private documents.
Figures
Forward citations
Cited by 1 Pith paper
-
Not All Entities are Created Equal: A Dynamic Anonymization Framework for Privacy-Preserving RAG
TRIP-RAG dynamically anonymizes only high-risk entities in RAG knowledge bases via three context-aware metrics, achieving privacy comparable to full anonymization with under 35% recall drop and up to 56% better genera...
Reference graph
Works this paper leans on
-
[1]
CJ Adams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017. Toxic Comment Classification Challenge. https: //kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge
2017
-
[2]
Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, and Siva Reddy. 2022. TopiOCQA: Open-domain Conversa- tional Question Answering with Topic Switching. Transactions of the Association for Computational Linguistics 10 (04 2022), 468–483. https://doi.org/10.1162/tacl_a_00471 arXiv:https://direct.mit.edu/tacl/article- pdf/doi/10.1162/tacl_...
-
[3]
Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021. Open-Domain Question Answering Goes Conversational via Question Rewriting. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2021)
2021
-
[4]
Simran Arora, Patrick Lewis, Angela Fan, Jacob Kahn, and Christopher Ré. 2023. Reasoning over Public and Private Data in Retrieval-Based Systems.Transactions of the Association for Computational Linguistics (2023). https://aclanthology.org/ 2023.tacl-1.51/
2023
-
[5]
Orlando Ayala and Patrice Bechard. 2024. Reducing hallucination in structured outputs via Retrieval-Augmented Generation. In Proceedings of the 2024 Con- ference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 6: Industry Track) , Yi Yang, Aida Davani, Avi Sil, and Anoop Kumar (Eds.). A...
-
[6]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https://arxiv.org/abs/1611.09268
arXiv 2018
-
[7]
Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, and Eneko Agirre. 2020. DoQA – Accessing Domain-Specific FAQs via Conversational QA. arXiv:2005.01328 [cs.CL] https://arxiv.org/abs/2005.01328
work page Pith review arXiv 2020
-
[8]
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Mike McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avirup Sil, Rosario Uceda-Sosa, Todd Ward, and Rong Zhang. 2019. The TechQA Dataset. arXiv:1911....
arXiv 2019
Show all 90 references
-
[9]
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceedings of the 2021 Conference on Empir...
2021
-
[10]
Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022. ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr...
2022 doi
-
[11]
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question Answering in Context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Ho...
2018 doi
-
[12]
Mark Collier and Joeran Beel. 2019. Memory-Augmented Neural Networks for Machine Translation. In Proceedings of Machine Translation Summit XVII: Research Track, Mikel Forcada, Andy Way, Barry Haddow, and Rico Sennrich (Eds.). European Association for Machine Translation, Dubli...
2019
-
[13]
Together Computer. 2023. RedPajama: an Open Dataset for Training Large Lan- guage Models. https://github.com/togethercomputer/RedPajama-Data
2023
-
[14]
Vinay Deolalikar. 2014. Distance or Coverage? Retrieving Knowledge-Rich Documents From Enterprise Text Collections. In Proceedings of the 23rd ACM In- ternational Conference on Conference on Information and Knowledge Management (Shanghai, China) (CIKM ’14). Association for Com...
2014
-
[15]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...
2024 arXiv
-
[16]
Ugur Guney, Volkan Cirik, and Kyunghyun Cho
Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine. arXiv:1704.05179 [cs.CL] https://arxiv.org/abs/1704.05179
2017 arXiv
-
[17]
Enron Corp and William W. Cohen. 2015. Enron Email Dataset. https: //www.loc.gov/item/2018487913/ United States Federal Energy Regulatory Com- mission, William W. Cohen, MLD, CMU, Philadelphia, PA. [Software, E-Resource]. Retrieved from the Library of Congress
2015
-
[18]
Song Feng, Hui Wan, Chulaka Gunasekara, Siva Patel, Sachindra Joshi, and Luis Lastras. 2020. doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor ...
2020 doi
-
[19]
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv preprint arXiv:2101.00027 (2020)
2020 arXiv
-
[20]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https://arxiv.org/abs/2312.10997
2024 arXiv
-
[21]
Samira Ghodratnama and Mehrdad Zakershahrak. 2024. Adapting LLMs for Effi- cient, Personalized Information Retrieval: Methods and Implications. In Service- Oriented Computing – ICSOC 2023 Workshops , Flavia Monti, Pierluigi Plebani, Naouel Moha, Hye-young Paik, Johanna Barzen,...
2024
-
[22]
Grand View Research. 2024. Retrieval Augmented Generation Market Size Report, 2030. https://www.grandviewresearch.com/industry-analysis/retrieval- augmented-generation-rag-market-report
2024
-
[23]
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014. Neural Turing Machines. arXiv:1410.5401 [cs.NE] https://arxiv.org/abs/1410.5401
2014 arXiv
-
[24]
Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort
Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L. Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort. 2023. Contin- ual Pre-Training of Large Language Models: How to (re)warm your model? arXiv:2308.04014 [cs.CL] https://arxiv.org/abs/2308.04014
2023 arXiv
-
[25]
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021. CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. arXiv:2103.06268 [cs.CL] https://arxiv.org/abs/2103.06268
2021 arXiv
-
[26]
Jing Huang, Diyi Yang, and Christopher Potts. 2024. Demystifying Verbatim Memorization in Large Language Models. arXiv:2407.17817 [cs.CL] https: //arxiv.org/abs/2407.17817
2024 arXiv
-
[27]
Yulong Hui, Yao Lu, and Huanchen Zhang. 2024. UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis. arXiv:2406.15187 [cs.AI] https://arxiv.org/abs/2406.15187
2024 arXiv
-
[28]
Infiniflow. 2024. RAGFlow: An open-source RAG (Retrieval-Augmented Gen- eration) engine based on deep document understanding. https://github.com/ infiniflow/ragflow Accessed: 2024-09-18
2024
-
[29]
Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based Neural Structured Learning for Sequential Question Answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Ling...
2017 doi
-
[30]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...
2024 arXiv
-
[31]
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu
-
[32]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehen- sion. arXiv e-prints, Article arXiv:1705.03551 (2017), arXiv:1705.03551 pages. arXiv:1705.03551 EnronQA: Towards Personaliz...
2017 arXiv
-
[33]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehen- sion. In Proceedings of the 55th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers) ...
2017 doi
-
[34]
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016. Bag of Tricks for Efficient Text Classification. arXiv preprint arXiv:1607.01759 (2016)
2016 arXiv
-
[35]
Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023. HAGRID: A Human-LLM Collaborative Dataset for Generative Information- Seeking with Attribution. arXiv:2307.16883 (2023)
2023 arXiv
-
[36]
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raf- fel. 2023. Large Language Models Struggle to Learn Long-Tail Knowledge. arXiv:2211.08411 [cs.CL] https://arxiv.org/abs/2211.08411
2023 arXiv
-
[37]
Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, and Bing Liu
-
[38]
Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav San- thanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023. DSPy: Com- piling Declarative Language Model Calls into ...
2023 arXiv
-
[39]
Bryan Klimt and Yiming Yang. 2004. The Enron Corpus: A New Dataset for Email Classification Research. In European Conference on Machine Learning . Springer Berlin Heidelberg, Berlin, Heidelberg, 217–226
2004
-
[40]
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Her- mann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA Reading Comprehension Challenge. Transactions of the Association for Computational Linguistics 6 (2018), 317–328. https://doi.org/10.11...
2018 doi
-
[41]
Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019
-
[42]
Jure Leskovec, Anand Rajaraman, and Jeffrey David Ullman. 2020. Mining of Massive Data Sets. Cambridge University Press
2020
-
[43]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.1140...
2021 arXiv
-
[44]
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Infor- mation Retrieval Research with Sparse and Dense Representations. InProceedings of the 44th Annual International ACM SIGIR Con...
2021
-
[45]
Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu. 2024. Rethinking Machine Unlearning for Large Language Models. arXiv:2402.08787 ...
2024 arXiv
-
[46]
Yougang Lyu, Lingyong Yan, Shuaiqiang Wang, Haibo Shi, Dawei Yin, Pengjie Ren, Zhumin Chen, Maarten de Rijke, and Zhaochun Ren. 2024. KnowTun- ing: Knowledge-aware Fine-tuning for Large Language Models. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural La...
2024 doi
-
[47]
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query Rewriting in Retrieval-Augmented Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Associ...
2023 doi
-
[48]
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary Chase Lipton, and J Zico Kolter. 2024. TOFU: A Task of Fictitious Unlearning for LLMs. InFirst Conference on Language Modeling. https://openreview.net/forum?id=B41hNBoWLo
2024
-
[49]
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024. ExpertQA: Expert-Curated Questions and Attributed Answers. arXiv:2309.07852 [cs.CL] https://arxiv.org/abs/2309.07852
2024 arXiv
-
[50]
Luke Merrick. 2024. Embedding And Clustering Your Data Can Improve Con- trastive Pretraining. arXiv:2407.18887 [cs.LG] https://arxiv.org/abs/2407.18887
2024 arXiv
-
[51]
Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020. COVID-QA: A Question Answering Dataset for COVID-19. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020 , Karin Verspoor, Kevin Bretonnel Cohen, Mark Dredze, Emilio Ferrara, Jonathan May, ...
2020
-
[52]
Chenghao Mou, Chris Ha, Kenneth Enevoldsen, and Peiyuan Liu. 2023. ChenghaoMou/text-dedup: Reference Snapshot. https://doi.org/10.5281/zenodo. 8364980
2023 doi
-
[53]
Kai Nakamura, Sharon Levy, Yi-Lin Tuan, Wenhu Chen, and William Yang Wang
-
[54]
Abhilash Nandy, Soumya Sharma, Shubham Maddhashiya, Kapil Sachdeva, Pawan Goyal, and NIloy Ganguly. 2021. Question Answering over Electronic Devices: A New Benchmark Dataset and a Multi-Task Learning based QA Frame- work. In Findings of the Association for Computational Lingui...
2021 doi
-
[55]
Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christo- pher Potts, Matei Zaharia, and Omar Khattab. 2024. Optimizing In- structions and Demonstrations for Multi-Stage Language Model Programs. arXiv:2406.11695 [cs.CL] https://arxiv.org/abs/2406.11695
2024 arXiv
-
[56]
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Mail- lard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Ta...
2021
-
[57]
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Join...
2019 doi
-
[58]
Fabio Pinelli, Gabriele Tolomei, and Giovanni Trappolini. 2023. FLIRT: Federated Learning for Information Retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for...
2023
-
[59]
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...
2022 arXiv
-
[60]
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. arXiv:1806.03822 [cs.CL] https: //arxiv.org/abs/1806.03822
2018 arXiv
-
[61]
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A Con- versational Question Answering Challenge. Transactions of the Association for Computational Linguistics 7 (2019), 249–266. https://doi.org/10.1162/tacl_a_00266
2019 doi
-
[62]
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How Much Knowledge Can You Pack Into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu ...
2020 doi
-
[63]
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh Menon, Md Parvez, and Zhe Feng. 2023. DelucionQA: Detecting Hallucinations in Domain-specific Question Answering. In Findings of the Association for Computational Linguistics: EMNLP 20...
2023 doi
-
[64]
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational ...
2022
-
[65]
Milad Shokouhi and Luo Si. 2011. Federated Search. Foundations and Trends® in Information Retrieval 5, 1 (2011), 1–102. https://doi.org/10.1561/1500000010
2011 doi
-
[66]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. arXiv preprint arXiv:2104.07567 (2021). https://arxiv.org/abs/2104.07567
2021 arXiv
-
[67]
Peters, Abhilasha Ravichander, Kyle Richardson, Ze- jiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkin- son, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas ...
2024 arXiv
-
[69]
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher Manning, and Chelsea Finn. 2023. Fine-tuning Language Models for Factuality. InNeurIPS 2023 Workshop on Instruction Tuning and Instruction Following . https://openreview.net/forum? id=kEK08VdSO5
2023
-
[70]
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. NewsQA: A Machine Comprehen- sion Dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP, Phil Blunsom, Antoine Bordes, Kyunghyun Cho, S...
2017 doi
-
[71]
Shuai Wang, Ekaterina Khramtsova, Shengyao Zhuang, and Guido Zuccon. 2024. FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval ...
2024
-
[72]
Wang and Duen Horng Chau
Zijie J. Wang and Duen Horng Chau. 2024. MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24) ....
2024
-
[73]
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. 2024. Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates. arXiv:2408.13006 [cs.CL] https://arxiv. org/abs/2408.13006
2024 arXiv
-
[74]
Zeqiu Wu, Ryu Parish, Hao Cheng, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. InSCIt: Information-Seeking Conver- sations with Mixed-Initiative Interactions.Transactions of the Association for Com- putational Linguistics 11 (2023), 453–468....
2023 doi
-
[75]
Sirui Xia, Xintao Wang, Jiaqing Liang, Yifei Zhang, Weikang Zhou, Jiaji Deng, Fei Yu, and Yanghua Xiao. 2024. Ground Every Sentence: Improv- ing Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation. arXiv:2407.01796 [cs.CL] https://arxiv.org/abs/2407.01796
2024 arXiv
-
[76]
Zitong Yang, Neil Band, Shuangping Li, Emmanuel Candes, and Tatsunori Hashimoto. 2025. Synthetic continued pretraining. In The Thirteenth Interna- tional Conference on Learning Representations . https://openreview.net/forum? id=07yvxWDSla
2025
-
[77]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...
2018
-
[78]
Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2024. Mitigating the Pri- vacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data. arXiv:2406.14773 [cs.CR] https://arxiv.org/abs/2406.14773
2024 arXiv
-
[79]
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, and Jiliang Tang. 2024. The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG). arXiv:2402.16893 [cs.CR] https://arxiv.org/abs...
2024 arXiv
-
[80]
Saber Zerhoudi and Michael Granitzer. 2024. PersonaRAG: Enhanc- ing Retrieval-Augmented Generation Systems with User-Centric Agents. arXiv:2407.09394 [cs.IR] https://arxiv.org/abs/2407.09394
2024
-
[81]
Asm Dem Plan
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A Question An- swering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In Proceedings of the 59th Annual Meeting of the Association for ...
2021
-
[86]
the email
The question should be specific to the factual contents of this email. In other words, if you were given 100 emails, would you know that the question was about this email and not another one? Phrases such as "the email" or "the message" are not specific enough unless grounded ...
-
[87]
It is okay to use the sender and recipients as context, but the question should not be about them
The question should focus on the main contents of the message, not on the formatting, the sender, or recipients. It is okay to use the sender and recipients as context, but the question should not be about them. It is okay to ask about things like the cell phone numbers or con...
-
[88]
It should not be a matter of opinion or require any interpretation
The question should be objective and answerable with a single sentence. It should not be a matter of opinion or require any interpretation
-
[89]
The question should be realistic to what a person might ask about an email they received, especially in the context of working in a professional setting and recalling important details
-
[90]
The question should NOT require any external knowledge beyond the contents of the email itself
-
[91]
How many times did the sender mention the word ’important’?
The question should NOT include counting, math, or any other operations. For example "How many times did the sender mention the word ’important’?" or "how many recipients were there?" are not allowed. It is okay for the question to ask about a number in the email such as a per...
-
[2019]
PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vince...
2019 doi
-
[2022]
In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.)
HybriDialogue: An Information-Seeking Dialogue Dataset Grounded on Tabular and Textual Data. In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dubl...
2022 doi
-
[2023]
In The Eleventh International Conference on Learning Representations
Continual Pre-training of Language Models. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=m_ GDIItaI3o
-
[3734]
https://doi.org/10.18653/v1/2022.naacl-main.272
2022 doi
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.