Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

EnronQA: Towards Personalized RAG over Private Documents

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read EnronQA, a benchmark of 103,638 private emails and 528,304 question-answer pairs across 150 inboxes, is calibrated so end-to-end RAG accuracy tracks retrieval quality rather than LLM memorization, with no-context accuracy below 5% and…

desk verdict Large, useful private-document RAG benchmark; the calibration claim is partly self-enforcing and human validation is thin, but the resource deserves review. read the letter →

arxiv 2505.00263 v1 pith:MG4SEZAH submitted 2025-05-01 cs.IR cs.CL

classification cs.IRcs.CL
keywords EnronQAretrieval-augmentedgenerationprivate-documentbenchmarkquestion-answerretrievalcalibrationpersonalizedmemorizationLoRA
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

EnronQA turns the cleaned Enron email corpus into a benchmark for retrieval-augmented generation over private documents: 103,638 emails, 528,304 synthetic question-answer pairs, and 150 separate user inboxes. The paper's central claim is that this benchmark is properly calibrated: without any retrieved context, a large language model scores below 5%, and every additional point of retrieval recall adds roughly 0.6 points of answer accuracy. That contrasts with Wikipedia-based benchmarks such as NaturalQuestions and TriviaQA, where a model can often answer from memorized knowledge alone and a retriever must be quite good before it helps at all. If the calibration claim holds, EnronQA lets practitioners read end-to-end RAG accuracy as a direct measure of retriever quality, and it provides a test bed for personalized retrieval over realistic per-user document stores.

What carries the argument

The paper's working parts are the four-gate question admission pipeline and the per-inbox segmentation of the corpus. A question enters the benchmark only if it is specific, objective, grounded, and high quality by automated checks; this is what keeps no-context accuracy near zero and makes end-to-end accuracy move with retrieval recall. The 150 distinct inboxes are what let the same corpus support personalized retrieval experiments rather than only global retrieval.

What would settle it

Have independent human annotators answer a random sample of roughly 500 EnronQA questions with no email in context; if more than a small percentage are answered correctly, the groundedness check is too lenient and the benchmark would not track retrieval quality. The same test can be rerun on each new generation of large language models, since the no-context baseline rising above 5% would signal that the corpus has leaked into training.

Watch

Extended reading notes

Core claim

EnronQA is constructed by filtering the cleaned Enron email corpus down to 103,638 usable emails and generating 528,304 QA pairs with a multi-stage LLM pipeline. Each question must pass four automated checks before entering the benchmark: specificity, objectivity, groundedness, and a rules-based quality gate. The central discovery is about calibration: in simulated Recall@1 sweeps, EnronQA is the only benchmark tested where adding the correct context always improves accuracy over the no-context baseline, with no-context accuracy below 5% and nearly a 0.6% accuracy gain per 1% recall gain. The paper also reports baseline RAG results and a memorization case study in which LoRA adapters trained on up to 20,000 facts match long-context performance, while retrieval still outperforms both.

Load-bearing premise

The load-bearing premise is that the four automated checks performed by large language models correctly certify question quality, so the 528,304 question-answer pairs are genuinely answerable from their paired email and not from general knowledge.

Editorial extensions

If this is right

  • End-to-end accuracy on EnronQA can be read as a direct measure of retriever quality, since every point of recall adds roughly 0.6 points of accuracy.
  • Private-document RAG pipelines can be benchmarked without first controlling for whether the LLM already memorized the documents, which is not possible on Wikipedia-based benchmarks.
  • The benchmark's 150-inbox structure enables personalized retrieval evaluation, where systems must find and reason over the right user's documents.
  • The memorization case study shows LoRA adapters can recall up to 20,000 facts at a level comparable to putting the facts in context, while retrieval still outperforms both.
  • The benchmark is large enough for fine-tuning and continued-pretraining experiments, not just for evaluation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper leaves implicit: the specificity gate likely pushes questions toward distinctive named entities, which would explain the strong lexical-retrieval results; re-generating questions with paraphrased, entity-light wording would test whether dense retrieval closes the gap.
  • The calibration claim is time-sensitive; as future large language models train on broader data, the Enron corpus may leak into training, so re-running the no-context baseline on each new model generation is a simple contamination monitor.
  • The memorization results suggest an open capacity-versus-cost tradeoff; because LoRA adapters hold up to 20,000 facts, a natural next experiment is measuring at what fact count a large-rank adapter becomes cheaper or faster than maintaining a search index.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces EnronQA, a synthetic question-answering benchmark built from the cleaned Enron email corpus, containing 103,638 emails and 528,304 QA pairs across 150 user inboxes. The dataset is constructed through a DSPy-based pipeline that generates questions with Llama-3.1-70B and filters them through four LLM-based checks: specificity, objectivity, groundedness, and a rule-based quality judge. The authors present dataset statistics, a calibration study comparing EnronQA to NaturalQuestions and TriviaQA, baseline RAG experiments with BM25 and ColBERTv2 over several LLMs, and a case study on LoRA-based memorization of facts. The paper's central claims are that EnronQA is the only evaluated benchmark where adding retrieved context always improves over a no-context baseline, and that it provides a better-calibrated testbed for retriever quality than Wikipedia-based benchmarks.

Significance. If the quality of the 528k synthetic QA pairs is adequately validated, EnronQA would be a valuable community resource: it is large, segmented by user, drawn from realistic private emails, and released with prompts and artifacts. The calibration result, if it holds under independent evaluation, would be useful for the RAG community because it addresses the memorization problem that afflicts Wikipedia-based benchmarks. The LoRA memorization study is also a relevant initial exploration. The paper is transparent about its pipeline and releases substantial supporting material, including prompts and chain-of-thought outputs, which strengthens reproducibility. The main weaknesses are that the evidence for question quality is thin (41 human-labeled-quality examples and 200 judge agreement examples), the evaluation stack is heavily concentrated in one model family, and the headline calibration claim is partly enforced by construction rather than observed independently.

major comments (4)
  1. [§4.1, Table 3] Table 3 reports an average of 491.81 emails per user, but 103,638 emails divided by 150 users equals 690.92. This arithmetic inconsistency means the summary statistics as reported cannot be trusted until the table is corrected or the discrepancy is explained. Since the size and composition of the benchmark are central to the paper's claims, this needs to be fixed before publication.
  2. [§4.2 vs. §3.2] The headline calibration claim that 'EnronQA is the only benchmark where adding context is always better than the no-context baseline' is partly enforced by the groundedness filter in §3.2: a question is admitted only when both Llama-3.1-70B and Mixtral-8x7B fail to reproduce the gold answer without the email. The no-context accuracy below 5% is therefore a property of the benchmark's construction, not an independent discovery about private documents. The comparison to NaturalQuestions and TriviaQA, which were not filtered this way, should be framed accordingly, and the authors should report calibration on an unfiltered or differently filtered subset to support the stronger interpretation.
  3. [§3.2, §4.2, §5.2] The validation of question quality is too thin for a resource of 528k pairs. The LLM judge's 0.98 F1 was measured on only 200 generated match/no-match judgements, and this measures answer-equivalence, not whether a gold answer is factually correct or whether a question is answerable. The quality ruleset was tuned on 21 development questions and tested on 20, with an inter-annotator Spearman correlation of only 0.5. Because the same Llama-3.1-70B family generates questions, filters them, and scores final accuracy, there is a real risk of self-consistency bias. The authors should provide independent human evaluation on a larger random sample, report agreement statistics, and check for unanswerable or ambiguous questions before presenting EnronQA as 'high quality.'
  4. [§6, Table 5] The claim that LoRA memorization 'matches long-context performance at almost all scales' is not well supported by Table 5. At 5,000 facts the rank-1024 adapter collapses to 0.03 accuracy, and at 20,000 facts ranks 512, 1024, and 2048 yield 0.08, 0.00, and 0.03, respectively. Several cells show non-monotonic behavior across ranks, and no error bars, multiple seeds, or standard deviations are reported. This makes the conclusion about 'surprising capacity' fragile and should be either supported with variance estimates or substantially tempered.
minor comments (6)
  1. [§1] The text contains an unresolved placeholder '(??)' in the sentence about headroom for improving retriever quality; please replace it with the intended citation or number.
  2. [Fig. 2 caption] The caption refers to 'Mixtral-7B-Instruct' while the text and reference correctly identify the model as Mixtral-8x7B-Instruct; please correct the model name.
  3. [§5.2] There is a typo 'GPTo' that should read 'GPT-4o' or 'GPT4o' consistently.
  4. [Table 4] The column headers in Table 4 are difficult to parse because 'R@5' appears only under 'Query Rewrite' and the row labels mix retriever names with recall values. Please restructure the table so that each retriever's Recall@5 is clearly associated with the right column.
  5. [Appendix B.7] The introductory sentence in B.7 is duplicated from B.5 and says 'These prompts are used to both answer the question given the context of an email or to produce an answer to the question with no grounding,' which does not describe the quality-evaluation prompt in B.7; please correct it.
  6. [§6.1] The memorization experiment reports '10 epochs with rate 1e-4' but does not specify the optimizer, learning-rate schedule, or how accuracy is computed from the LLM judge's outputs; please add these details.

Circularity Check

1 steps flagged · score 4.0 of 10

Calibration claim is partly enforced by the groundedness filter: questions answerable without context are removed by construction, so the low no-context baseline is a selection effect rather than an independent discovery.

  1. self definitional [§3.2 QA Generation Pipeline (Groundedness); §4.2 Calibration]
    "Groundedness. We determine a question to be "grounded" if neither the llama nor mixtral model can answer the same as the gold answer given no email as context. This both tests that the answers to our questions are not memorized and that the questions are not easily guessable. ... We find that EnronQA is the only benchmark where adding context is always better than the no-context baseline."

    The groundedness criterion is an inclusion rule: a QA pair enters EnronQA only when both Llama-3.1-70B and Mixtral-8x7B, given no email, fail to produce the gold answer. The calibration section then measures no-context accuracy on exactly this filtered set and reports it is below 5% and that context always helps. That outcome is logically forced by the filter: any question the same judge could answer without context was discarded. Calling this "the knowledge in EnronQA has not been memorized" restates the admission rule rather than reporting an independent discovery. The Wikipedia comparison is not apples-to-apples, because those benchmarks were not filtered to exclude memorized or guessable questions.

full rationale

EnronQA itself is a corpus-derived resource, not a derivation, and its construction is largely independent: the 103,638-email corpus, filtering steps, and the public Enron source provide external grounding. The self-citation to DSPy/MIPROv2 (Ref. [55], an author of the present paper) is methodological and not load-bearing for the benchmark's validity. The main circularity is narrower: the headline calibration result ('only benchmark where adding context is always better', no-context below 5%) is partly enforced by the groundedness filter in §3.2, which removes questions that the same LLM family can answer without context. The benchmark can still be a useful resource, but the claimed advantage over Wikipedia benchmarks is a designed property, not a discovered one. The QA-quality evidence (0.98 F1 on 200 judge instances, ruleset on 41 labeled questions, same Llama-3.1-70B family for generation and final scoring) is thin but concerns validation strength, not derivation circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The benchmark's central claims rest on hand-set corpus filtering thresholds, LLM-judge validity, and the assumption that the Enron corpus mimics private enterprise documents. No new physical entities or fitted target results are introduced. The LoRA case study adds experiment-specific hyperparameters that are not part of the benchmark itself.

free parameters (3)
  • Corpus filtering thresholds = Jaccard 0.9, 9 bands, 27 rows; 50-1000 words; length 3-10; ratio >=0.65; ellipsis <10%; confidence 80-90%
    Hand-set cutoffs in Section 3.1 determine which emails enter the benchmark and therefore which questions can be generated; they are choices, not values derived from a target.
  • Specificity hard-negative count = k=10
    The specificity check retrieves 10 similar emails; the criterion for whether a question distinguishes the gold email is implicit in this choice.
  • LoRA training hyperparameters = ranks 8-2048; 10 epochs; lr 1e-4; alpha=4x rank; dropout 0.05
    Used in the memorization case study in Section 6.1; results vary strongly with rank and the setup is borrowed from TOFU.
assumptions (5)
  • domain assumption The 2015 Enron release is a realistic proxy for private enterprise email documents.
    The benchmark's usefulness for private RAG rests on this; Enron is a public corpus and older than modern enterprise mail, but is used as the stand-in in Sections 1 and 3.1.
  • domain assumption The LLM judge correctly determines answer equivalence.
    F1 0.98 on 200 samples supports it, but this is not a formal guarantee; the judge is used for dataset filtering and all accuracy scores.
  • ad hoc to paper The four quality criteria (specificity, objectivity, groundedness, quality) as implemented cover the relevant notion of question quality.
    These criteria are introduced in Section 3.2 and are operationalized with a rule set calibrated on 41 questions; no external benchmark validates them.
  • domain assumption No-context failure by Llama and Mixtral implies the answer is not in the parametric knowledge of the LLMs tested.
    Groundedness filtering in Section 3.2 assumes the two tested models are representative; a future model might memorize the Enron corpus and still answer.
  • domain assumption DSPy optimization and synthetic generation produce a diverse question distribution.
    There is no distributional analysis against real user queries; the benchmark may emphasize lexically specific entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of EnronQA: Towards Personalized RAG over Private Documents." pith.science (2026). https://pith.science/paper/MG4SEZAH

@misc{pith2026250500263,
  author       = {Pith},
  title        = {Pith review of: EnronQA: Towards Personalized RAG over Private Documents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MG4SEZAH}},
  note         = {Machine review of arXiv:2505.00263}
}
read the original abstract

Retrieval Augmented Generation (RAG) has become one of the most popular methods for bringing knowledge-intensive context to large language models (LLM) because of its ability to bring local context at inference time without the cost or data leakage risks associated with fine-tuning. A clear separation of private information from the LLM training has made RAG the basis for many enterprise LLM workloads as it allows the company to augment LLM's understanding using customers' private documents. Despite its popularity for private documents in enterprise deployments, current RAG benchmarks for validating and optimizing RAG pipelines draw their corpora from public data such as Wikipedia or generic web pages and offer little to no personal context. Seeking to empower more personal and private RAG we release the EnronQA benchmark, a dataset of 103,638 emails with 528,304 question-answer pairs across 150 different user inboxes. EnronQA enables better benchmarking of RAG pipelines over private data and allows for experimentation on the introduction of personalized retrieval settings over realistic data. Finally, we use EnronQA to explore the tradeoff in memorization and retrieval when reasoning over private documents.

Figures

Figures reproduced from arXiv: 2505.00263 by the authors.

Figure 1
Figure 1. The EnronQA benchmark enables personalized and private retrieval benchmarking on a cleaned corpus of over 100,000 emails spanning 528,304 quality question-answer pairs over 150 users. We explore both single and multi-user retrieval settings. Abstract Retrieval Augmented Generation (RAG) has become one of the most popular methods for bringing knowledge-intensive context to large language models (LLM) because of its a… view at source ↗
Figure 2
Figure 2. Our multistage compound LLM system for QA Generation on the Enron emails corpus. Our pipeline consists of 4 stages [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Question Rewrite Pipeline. First, we ask Llama 3.1 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Calibration experiment results. Although all benchmarks scale roughly linearly with more accurate context, [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not All Entities are Created Equal: A Dynamic Anonymization Framework for Privacy-Preserving RAG

    cs.CR 2026-03 unverdicted novelty 7.0 of 10

    TRIP-RAG dynamically anonymizes only high-risk entities in RAG knowledge bases via three context-aware metrics, achieving privacy comparable to full anonymization with under 35% recall drop and up to 56% better genera...

Reference graph

Works this paper leans on

90 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    CJ Adams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017. Toxic Comment Classification Challenge. https: //kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge

  2. [2]

    Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, and Siva Reddy. 2022. TopiOCQA: Open-domain Conversa- tional Question Answering with Topic Switching. Transactions of the Association for Computational Linguistics 10 (04 2022), 468–483. https://doi.org/10.1162/tacl_a_00471 arXiv:https://direct.mit.edu/tacl/article- pdf/doi/10.1162/tacl_...

  3. [3]

    Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021. Open-Domain Question Answering Goes Conversational via Question Rewriting. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2021)

  4. [4]

    Simran Arora, Patrick Lewis, Angela Fan, Jacob Kahn, and Christopher Ré. 2023. Reasoning over Public and Private Data in Retrieval-Based Systems.Transactions of the Association for Computational Linguistics (2023). https://aclanthology.org/ 2023.tacl-1.51/

  5. [5]

    Orlando Ayala and Patrice Bechard. 2024. Reducing hallucination in structured outputs via Retrieval-Augmented Generation. In Proceedings of the 2024 Con- ference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 6: Industry Track) , Yi Yang, Aida Davani, Avi Sil, and Anoop Kumar (Eds.). A...

  6. [6]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https://arxiv.org/abs/1611.09268

  7. [7]

    Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, and Eneko Agirre. 2020. DoQA – Accessing Domain-Specific FAQs via Conversational QA. arXiv:2005.01328 [cs.CL] https://arxiv.org/abs/2005.01328

  8. [8]

    Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Mike McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avirup Sil, Rosario Uceda-Sosa, Todd Ward, and Rong Zhang. 2019. The TechQA Dataset. arXiv:1911....

Show all 90 references
  1. [9]

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceedings of the 2021 Conference on Empir...

  2. [10]

    Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022. ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr...

  3. [11]

    Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question Answering in Context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Ho...

  4. [12]

    Mark Collier and Joeran Beel. 2019. Memory-Augmented Neural Networks for Machine Translation. In Proceedings of Machine Translation Summit XVII: Research Track, Mikel Forcada, Andy Way, Barry Haddow, and Rico Sennrich (Eds.). European Association for Machine Translation, Dubli...

  5. [13]

    Together Computer. 2023. RedPajama: an Open Dataset for Training Large Lan- guage Models. https://github.com/togethercomputer/RedPajama-Data

  6. [14]

    Vinay Deolalikar. 2014. Distance or Coverage? Retrieving Knowledge-Rich Documents From Enterprise Text Collections. In Proceedings of the 23rd ACM In- ternational Conference on Conference on Information and Knowledge Management (Shanghai, China) (CIKM ’14). Association for Com...

  7. [15]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...

  8. [16]

    Ugur Guney, Volkan Cirik, and Kyunghyun Cho

    Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine. arXiv:1704.05179 [cs.CL] https://arxiv.org/abs/1704.05179

  9. [17]

    Enron Corp and William W. Cohen. 2015. Enron Email Dataset. https: //www.loc.gov/item/2018487913/ United States Federal Energy Regulatory Com- mission, William W. Cohen, MLD, CMU, Philadelphia, PA. [Software, E-Resource]. Retrieved from the Library of Congress

  10. [18]

    Song Feng, Hui Wan, Chulaka Gunasekara, Siva Patel, Sachindra Joshi, and Luis Lastras. 2020. doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor ...

  11. [19]

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv preprint arXiv:2101.00027 (2020)

  12. [20]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https://arxiv.org/abs/2312.10997

  13. [21]

    Samira Ghodratnama and Mehrdad Zakershahrak. 2024. Adapting LLMs for Effi- cient, Personalized Information Retrieval: Methods and Implications. In Service- Oriented Computing – ICSOC 2023 Workshops , Flavia Monti, Pierluigi Plebani, Naouel Moha, Hye-young Paik, Johanna Barzen,...

  14. [22]

    Grand View Research. 2024. Retrieval Augmented Generation Market Size Report, 2030. https://www.grandviewresearch.com/industry-analysis/retrieval- augmented-generation-rag-market-report

  15. [23]

    Alex Graves, Greg Wayne, and Ivo Danihelka. 2014. Neural Turing Machines. arXiv:1410.5401 [cs.NE] https://arxiv.org/abs/1410.5401

  16. [24]

    Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort

    Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L. Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort. 2023. Contin- ual Pre-Training of Large Language Models: How to (re)warm your model? arXiv:2308.04014 [cs.CL] https://arxiv.org/abs/2308.04014

  17. [25]

    Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021. CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. arXiv:2103.06268 [cs.CL] https://arxiv.org/abs/2103.06268

  18. [26]

    Jing Huang, Diyi Yang, and Christopher Potts. 2024. Demystifying Verbatim Memorization in Large Language Models. arXiv:2407.17817 [cs.CL] https: //arxiv.org/abs/2407.17817

  19. [27]

    Yulong Hui, Yao Lu, and Huanchen Zhang. 2024. UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis. arXiv:2406.15187 [cs.AI] https://arxiv.org/abs/2406.15187

  20. [28]

    Infiniflow. 2024. RAGFlow: An open-source RAG (Retrieval-Augmented Gen- eration) engine based on deep document understanding. https://github.com/ infiniflow/ragflow Accessed: 2024-09-18

  21. [29]

    Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based Neural Structured Learning for Sequential Question Answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Ling...

  22. [30]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  23. [31]

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu

  24. [32]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehen- sion. arXiv e-prints, Article arXiv:1705.03551 (2017), arXiv:1705.03551 pages. arXiv:1705.03551 EnronQA: Towards Personaliz...

  25. [33]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehen- sion. In Proceedings of the 55th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers) ...

  26. [34]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016. Bag of Tricks for Efficient Text Classification. arXiv preprint arXiv:1607.01759 (2016)

  27. [35]

    Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023. HAGRID: A Human-LLM Collaborative Dataset for Generative Information- Seeking with Attribution. arXiv:2307.16883 (2023)

  28. [36]

    Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raf- fel. 2023. Large Language Models Struggle to Learn Long-Tail Knowledge. arXiv:2211.08411 [cs.CL] https://arxiv.org/abs/2211.08411

  29. [37]

    Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, and Bing Liu

  30. [38]

    Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts

    Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav San- thanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023. DSPy: Com- piling Declarative Language Model Calls into ...

  31. [39]

    Bryan Klimt and Yiming Yang. 2004. The Enron Corpus: A New Dataset for Email Classification Research. In European Conference on Machine Learning . Springer Berlin Heidelberg, Berlin, Heidelberg, 217–226

  32. [40]

    Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Her- mann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA Reading Comprehension Challenge. Transactions of the Association for Computational Linguistics 6 (2018), 317–328. https://doi.org/10.11...

  33. [41]

    Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  34. [42]

    Jure Leskovec, Anand Rajaraman, and Jeffrey David Ullman. 2020. Mining of Massive Data Sets. Cambridge University Press

  35. [43]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.1140...

  36. [44]

    Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Infor- mation Retrieval Research with Sparse and Dense Representations. InProceedings of the 44th Annual International ACM SIGIR Con...

  37. [45]

    Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu

    Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu. 2024. Rethinking Machine Unlearning for Large Language Models. arXiv:2402.08787 ...

  38. [46]

    Yougang Lyu, Lingyong Yan, Shuaiqiang Wang, Haibo Shi, Dawei Yin, Pengjie Ren, Zhumin Chen, Maarten de Rijke, and Zhaochun Ren. 2024. KnowTun- ing: Knowledge-aware Fine-tuning for Large Language Models. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural La...

  39. [47]

    Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query Rewriting in Retrieval-Augmented Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Associ...

  40. [48]

    Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary Chase Lipton, and J Zico Kolter. 2024. TOFU: A Task of Fictitious Unlearning for LLMs. InFirst Conference on Language Modeling. https://openreview.net/forum?id=B41hNBoWLo

  41. [49]

    Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024. ExpertQA: Expert-Curated Questions and Attributed Answers. arXiv:2309.07852 [cs.CL] https://arxiv.org/abs/2309.07852

  42. [50]

    Luke Merrick. 2024. Embedding And Clustering Your Data Can Improve Con- trastive Pretraining. arXiv:2407.18887 [cs.LG] https://arxiv.org/abs/2407.18887

  43. [51]

    Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020. COVID-QA: A Question Answering Dataset for COVID-19. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020 , Karin Verspoor, Kevin Bretonnel Cohen, Mark Dredze, Emilio Ferrara, Jonathan May, ...

  44. [52]

    Chenghao Mou, Chris Ha, Kenneth Enevoldsen, and Peiyuan Liu. 2023. ChenghaoMou/text-dedup: Reference Snapshot. https://doi.org/10.5281/zenodo. 8364980

  45. [53]

    Kai Nakamura, Sharon Levy, Yi-Lin Tuan, Wenhu Chen, and William Yang Wang

  46. [54]

    Abhilash Nandy, Soumya Sharma, Shubham Maddhashiya, Kapil Sachdeva, Pawan Goyal, and NIloy Ganguly. 2021. Question Answering over Electronic Devices: A New Benchmark Dataset and a Multi-Task Learning based QA Frame- work. In Findings of the Association for Computational Lingui...

  47. [55]

    Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christo- pher Potts, Matei Zaharia, and Omar Khattab. 2024. Optimizing In- structions and Demonstrations for Multi-Stage Language Model Programs. arXiv:2406.11695 [cs.CL] https://arxiv.org/abs/2406.11695

  48. [56]

    Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Mail- lard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Ta...

  49. [57]

    Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Join...

  50. [58]

    Fabio Pinelli, Gabriele Tolomei, and Giovanni Trappolini. 2023. FLIRT: Federated Learning for Information Retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for...

  51. [59]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...

  52. [60]

    Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. arXiv:1806.03822 [cs.CL] https: //arxiv.org/abs/1806.03822

  53. [61]

    Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A Con- versational Question Answering Challenge. Transactions of the Association for Computational Linguistics 7 (2019), 249–266. https://doi.org/10.1162/tacl_a_00266

  54. [62]

    Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How Much Knowledge Can You Pack Into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu ...

  55. [63]

    Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh Menon, Md Parvez, and Zhe Feng. 2023. DelucionQA: Detecting Hallucinations in Domain-specific Question Answering. In Findings of the Association for Computational Linguistics: EMNLP 20...

  56. [64]

    Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational ...

  57. [65]

    Milad Shokouhi and Luo Si. 2011. Federated Search. Foundations and Trends® in Information Retrieval 5, 1 (2011), 1–102. https://doi.org/10.1561/1500000010

  58. [66]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. arXiv preprint arXiv:2104.07567 (2021). https://arxiv.org/abs/2104.07567

  59. [67]

    Peters, Abhilasha Ravichander, Kyle Richardson, Ze- jiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkin- son, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas ...

  60. [69]

    Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher Manning, and Chelsea Finn. 2023. Fine-tuning Language Models for Factuality. InNeurIPS 2023 Workshop on Instruction Tuning and Instruction Following . https://openreview.net/forum? id=kEK08VdSO5

  61. [70]

    Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. NewsQA: A Machine Comprehen- sion Dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP, Phil Blunsom, Antoine Bordes, Kyunghyun Cho, S...

  62. [71]

    Shuai Wang, Ekaterina Khramtsova, Shengyao Zhuang, and Guido Zuccon. 2024. FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval ...

  63. [72]

    Wang and Duen Horng Chau

    Zijie J. Wang and Duen Horng Chau. 2024. MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24) ....

  64. [73]

    Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. 2024. Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates. arXiv:2408.13006 [cs.CL] https://arxiv. org/abs/2408.13006

  65. [74]

    Zeqiu Wu, Ryu Parish, Hao Cheng, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. InSCIt: Information-Seeking Conver- sations with Mixed-Initiative Interactions.Transactions of the Association for Com- putational Linguistics 11 (2023), 453–468....

  66. [75]

    Sirui Xia, Xintao Wang, Jiaqing Liang, Yifei Zhang, Weikang Zhou, Jiaji Deng, Fei Yu, and Yanghua Xiao. 2024. Ground Every Sentence: Improv- ing Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation. arXiv:2407.01796 [cs.CL] https://arxiv.org/abs/2407.01796

  67. [76]

    Zitong Yang, Neil Band, Shuangping Li, Emmanuel Candes, and Tatsunori Hashimoto. 2025. Synthetic continued pretraining. In The Thirteenth Interna- tional Conference on Learning Representations . https://openreview.net/forum? id=07yvxWDSla

  68. [77]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...

  69. [78]

    Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2024. Mitigating the Pri- vacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data. arXiv:2406.14773 [cs.CR] https://arxiv.org/abs/2406.14773

  70. [79]

    Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, and Jiliang Tang. 2024. The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG). arXiv:2402.16893 [cs.CR] https://arxiv.org/abs...

  71. [80]

    Saber Zerhoudi and Michael Granitzer. 2024. PersonaRAG: Enhanc- ing Retrieval-Augmented Generation Systems with User-Centric Agents. arXiv:2407.09394 [cs.IR] https://arxiv.org/abs/2407.09394

  72. [81]

    Asm Dem Plan

    Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A Question An- swering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In Proceedings of the 59th Annual Meeting of the Association for ...

  73. [86]

    the email

    The question should be specific to the factual contents of this email. In other words, if you were given 100 emails, would you know that the question was about this email and not another one? Phrases such as "the email" or "the message" are not specific enough unless grounded ...

  74. [87]

    It is okay to use the sender and recipients as context, but the question should not be about them

    The question should focus on the main contents of the message, not on the formatting, the sender, or recipients. It is okay to use the sender and recipients as context, but the question should not be about them. It is okay to ask about things like the cell phone numbers or con...

  75. [88]

    It should not be a matter of opinion or require any interpretation

    The question should be objective and answerable with a single sentence. It should not be a matter of opinion or require any interpretation

  76. [89]

    The question should be realistic to what a person might ask about an email they received, especially in the context of working in a professional setting and recalling important details

  77. [90]

    The question should NOT require any external knowledge beyond the contents of the email itself

  78. [91]

    How many times did the sender mention the word ’important’?

    The question should NOT include counting, math, or any other operations. For example "How many times did the sender mention the word ’important’?" or "how many recipients were there?" are not allowed. It is okay for the question to ask about a number in the email such as a per...

  79. [2019]

    PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vince...

  80. [2022]

    In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.)

    HybriDialogue: An Information-Seeking Dialogue Dataset Grounded on Tabular and Textual Data. In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dubl...

  81. [2023]

    In The Eleventh International Conference on Learning Representations

    Continual Pre-training of Language Models. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=m_ GDIItaI3o

  82. [3734]

    https://doi.org/10.18653/v1/2022.naacl-main.272

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.