REVIEW 3 major objections 5 minor 61 references
Optimizing Knowledge Integration in Retrieval-Augmented Generation with Self-Selection
T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read A retrieval-augmented generation system can be improved by making the LLM generate two candidate answers—one from its own memory, one from retrieved passages—and then train it with direct preference optimization to choose the better one.
desk verdict A sensible DPO-based self-selection method for RAG whose evaluation is too thin to support the 'consistent' claim; worth refereeing after stronger stats and artifact release. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the two-stage generate-then-select loop of Eqs. (4)-(6), in which the same LLM produces (answer, explanation) pairs from parametric knowledge alone and from retrieved passages, then consumes both pairs and chooses the final answer. The selection step is trained with Direct Preference Optimization (DPO) on the Retrieval-Generation Preference (RGP) dataset: 3,756 instances, each with a query, a golden answer, one correct candidate, and one incorrect candidate, each with an explanation, augmented by treating answers to similar queries as additional negatives. The explanations do the load-bearing work: they are the evidence the selection prompt uses to arbitrate between candidates, and DPO rewards choosing the explanation-answer pair that matches the golden answer.
What would settle it
Run the selection step on held-out pairs where exactly one candidate is correct, using the trained model with candidate explanations generated independently at inference time and with candidate order randomized; if selection accuracy on these pairs is statistically indistinguishable from chance, or if a rule that always prefers the RAG answer matches the trained model's accuracy, then the claimed gains come from generation rather than from learned selection.
Extended reading notes
Core claim
The paper's central claim is that knowledge integration in RAG should be framed as a preference-selection problem the LLM can solve, not a fusion problem solved by the retriever or a fixed rule. The Self-Selection-RGP training makes the LLM both a better generator and a better selector: DPO on the augmented RGP dataset teaches it to prefer the correct candidate over the incorrect one, and this transfers to improved answer generation even without retrieval. The reported result is consistent gains over Standard RAG and SURE in almost all settings, with the largest margins on TriviaQA.
Load-bearing premise
The method assumes the LLM's written explanations are faithful enough that a DPO-trained model can learn to pick the correct answer from them; if explanations are uninformative or subtly wrong, the selection step cannot work, and the paper's own error analysis attributes 14% of errors to reasoning and 12% to selection.
Editorial extensions
If this is right
- Self-Selection-RGP improves over Standard RAG and SURE on most settings, for example gaining 7.2 accuracy points over Standard RAG on TriviaQA zero-shot with Mistral-7B.
- Training on the RGP dataset improves both selection and generation: the trained model answers better even without retrieval on TriviaQA, and consistently better with retrieval.
- The method is robust to retrieval changes: gains persist when swapping BGE for BM25 and across 1 to 10 retrieved passages, with smaller zero-shot to few-shot variance than Standard RAG.
- Removing either dataset augmentation or DPO alignment degrades accuracy, with the removal of preference alignment causing the larger drop.
- The authors state that the 3,756-instance RGP dataset will be released for future research.
Reading between the lines
- Beyond the paper, the same generate-both-then-select recipe could generalize to more than two candidates or to other dual-source settings such as conflicting documents or multiple retrievers, which the paper does not test.
- Because the selection prompt reasons over explanations, the method's ceiling is tied to explanation faithfulness; improving explanation quality, for example by training on explanation-verified pairs, may yield larger gains than those reported.
- The paper's error analysis attributes 14% of errors to reasoning and 12% to selection, suggesting roughly a quarter of current errors could disappear if selection and reasoning were made stronger.
- A testable extension is whether RGP data generated by a stronger or weaker LLM changes the transfer to the open-source models used here, which would clarify how much of the gain comes from the training data generator.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Self-Selection-RGP, a retrieval-augmented generation framework in which an LLM generates two candidate answers—one from parametric knowledge alone and one from retrieved passages—and is then prompted to select the better of the two. To improve selection accuracy, the authors construct a preference dataset (RGP) from WebQuestions, SQuAD 2.0, and SciQ using GPT-3.5 for candidate generation and golden-answer-based labeling, augment it by treating answers to similar queries as additional negatives, and fine-tune open-source LLMs (Mistral-7B, Llama2-13B-Chat) with Direct Preference Optimization (DPO). Experiments on Natural Questions and TriviaQA compare Self-Selection-RGP against LLM-only, Standard RAG, Self-RAG, and SURE across zero-shot and few-shot settings, with additional studies on different retrievers, varying numbers of retrieved passages, and ablations. The central claim is that Self-Selection-RGP consistently and robustly improves RAG accuracy across retrieval settings and base LLMs.
Significance. If the central claim holds, the work contributes a simple and modular method for knowledge integration in RAG: instead of adding a separate verification module or complexity, it trains the LLM itself to arbitrate between parametric and retrieval-based answers. The method is reproducible in principle—the authors promise to release the 3,756-instance RGP dataset, and the training recipe (LoRA + DPO on an automatically built preference set) is standard. The framework is also conceptually clean, extending the recent line of work on adaptive retrieval with a preference-alignment twist. The claimed gains over SURE and Standard RAG on TriviaQA are nontrivial and the analysis of answer-generation capability (though flawed, see major comments) tries to separate selection from generation effects. However, the evaluation's statistical fragility and one notable failure case (Llama2 zero-shot NQ) currently limit the strength of the 'consistent' claim.
major comments (3)
- [Table 2, Section 3.2] The evaluation uses only 500 questions per dataset and reports no variance, confidence intervals, significance tests, or multiple seeds. With n=500, the standard error of an accuracy estimate is at most about 2.2 percentage points, so several reported gains (e.g., +2.6 for Mistral on NQ zero-shot, +1.2 for Llama2 few-shot NQ over SURE) are within roughly one standard error and cannot be distinguished from noise. The paper's headline claim of 'consistent' superiority is therefore not statistically supported by the presented numbers. I recommend either reporting variance over multiple evaluation subsets or seeds, or tempering the claim to 'often improves' with the observed exceptions clearly flagged.
- [Table 2, Llama2-13B-Chat zero-shot NQ row] In a headline configuration, Self-Selection-RGP scores 46.2 Acc versus SURE's 52.0, a 5.8-point deficit, while also underperforming Standard RAG (45.2) only slightly. The paper attributes this to NQ questions being harder for LLMs, citing LLM-only performance, but that attribution is not evidence and does not explain why the deficit is specific to Self-Selection-RGP in the zero-shot case while the few-shot setting shows an improvement. This single result directly contradicts the claim of 'consistently' high effectiveness and 'robustness and stability' across settings. The paper should either provide a mechanistic explanation backed by analysis or revise the central claim to acknowledge this failure mode.
- [Section 3.4, Figure 6] The analysis claims that preference-alignment training improves the LLM's inherent answer-generation ability, but the results in Figure 6(a) and (b) show that Self-Selection-RGP-7B is worse than Mistral-7B on NQ without retrieval in both zero-shot and few-shot settings. The paper acknowledges this within the paragraph but then concludes that training 'has led to notable improvements in LLMs' ability to generate high-quality answers.' This overgeneralization is not supported by the paper's own data; the conclusion should be restricted to the retrieval-augmented setting, where the improvements are consistent, or the authors should analyze why generation degrades on NQ without retrieval.
minor comments (5)
- [Abstract] The dataset name is spelled 'TrivialQA' in the abstract; it should be 'TriviaQA' as used in the rest of the paper.
- [Section 2.3.2, end of paragraph] There is a typo: 'reseach' should be 'research'.
- [Section 3.1.2, Self-RAG description] The paper excludes Self-RAG on NQ because its training data includes NQ, but does not state whether SURE's training or development sets also overlap with NQ/TriviaQA; given SURE is the strongest baseline in several settings, this should be clarified for a fair comparison.
- [Section 2.3.3, Eq. (10)] The augmentation assumes that answers to similar queries are always invalid for a given query, which can introduce false negatives. The paper reports the number of augmented instances (21,928) but does not analyze how many of these negative pairs might actually contain correct answers; a small validation of this assumption would strengthen the method's motivation.
- [Table 3] The error analysis categorizes only 100 errors from TriviaQA with Mistral-7B. It would be helpful to indicate whether this sample is representative and to report the error distribution for NQ or for Llama2, since the method's performance varies substantially across those settings.
Circularity Check
No circularity: the method trains on golden-labeled preference pairs from WebQuestions/SQuAD2.0/SciQ and evaluates on held-out NQ/TriviaQA with exact-match metrics.
full rationale
The paper's derivation chain is self-contained and does not reduce any prediction to its own inputs. The RGP preference dataset is constructed by generating candidate answers with GPT-3.5 and retaining only pairs where one answer matches the golden answer and the other does not (Eq. 9); the golden labels come from existing QA datasets, not from the method's own outputs. DPO training (Eq. 12) then optimizes the open-source LLM to prefer the positive answer, and evaluation is performed on held-out NQ and TriviaQA test splits (500 questions each) using Exact Match, F1, and Accuracy, which are standard external metrics. No fitted parameter is renamed as a prediction, and no self-citation or uniqueness-theorem argument is used to force the design choice. The acknowledged zero-shot Llama2-13B deficit on NQ (46.2 vs SURE 52.0) is a consistency/statistical concern about the headline claim, not a circularity. The reliance on GPT-3.5 to generate explanations and to filter preference pairs is a data-quality and faithfulness issue, but it does not make the evaluation equivalent to the training signal by construction. The ablation study compares against Standard RAG and removes augmentation or alignment, providing independent checks of the design components. Overall, no circular step can be exhibited from the paper's equations or citations.
Assumptions & free parameters
free parameters (4)
- Retrieval top-K for inference =
5 (varied 1-10)
- Augmentation top-K (similar queries) =
not reported
- DPO beta =
not reported
- LoRA rank and alpha =
not reported
assumptions (5)
- standard math DPO objective (Eq. 12) is a valid and applicable optimization for preference alignment.
- domain assumption The BGE and BM25 retrievers surface passages that support the correct answer for a sufficient fraction of queries.
- domain assumption GPT-3.5 labeling (correct vs. incorrect against the golden answer) is reliable enough to build the preference pairs.
- ad hoc to paper Answers with similar queries are valid negative responses for a given query.
- domain assumption The explanations generated alongside each answer are faithful enough for the model to use them in selection.
Cite this review
Pith. "Pith review of Optimizing Knowledge Integration in Retrieval-Augmented Generation with Self-Selection." pith.science (2026). https://pith.science/paper/2XWK37I5
@misc{pith2026250206148,
author = {Pith},
title = {Pith review of: Optimizing Knowledge Integration in Retrieval-Augmented Generation with Self-Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/2XWK37I5}},
note = {Machine review of arXiv:2502.06148}
}
read the original abstract
Retrieval-Augmented Generation (RAG), which integrates external knowledge into Large Language Models (LLMs), has proven effective in enabling LLMs to produce more accurate and reliable responses. However, it remains a significant challenge how to effectively integrate external retrieved knowledge with internal parametric knowledge in LLMs. In this work, we propose a novel Self-Selection RAG framework, where the LLM is made to select from pairwise responses generated with internal parametric knowledge solely and with external retrieved knowledge together to achieve enhanced accuracy. To this end, we devise a Self-Selection-RGP method to enhance the capabilities of the LLM in both generating and selecting the correct answer, by training the LLM with Direct Preference Optimization (DPO) over a curated Retrieval Generation Preference (RGP) dataset. Experimental results with two open-source LLMs (i.e., Llama2-13B-Chat and Mistral-7B) well demonstrate the superiority of our approach over other baseline methods on Natural Questions (NQ) and TrivialQA datasets.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen. 2023. Retrieval-based Language Models and Applications. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 6: Tutorial Abstracts) . Association for Computational Linguistics, 41–46. https://aclanthology.org/2023. acl-tutorials.6
work page 2023
-
[2]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi
-
[3]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...
arXiv 2022
-
[4]
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic Parsing on Freebase from Question-Answer Pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 1533–1544
work page 2013
-
[5]
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. InPars: Unsupervised Dataset Generation for Information Retrieval. In Proceed- ings of the 45th International ACM SIGIR Conference on Research and Develop- ment in Information Retrieval (SIGIR ’22) . Association for Computing Machinery, 2387–2392
work page 2022
-
[6]
Table 3: Examples of errors and corresponding percentages
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
work page 2018
-
[7]
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Proceedings of the 31st International Conference on Neural Information Processing Systems. Curran Associates Inc., 4302–4310
work page 2017
-
[8]
Guanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang, Zhicheng Dou, and Ji-Rong Wen. 2024. Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation. arXiv:2406.18676
arXiv 2024
Show all 61 references
-
[9]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang
-
[10]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . https: //openreview.net/forum?id=nZeVKeeFYf9
2022
-
[11]
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave
-
[12]
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park
-
[13]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023
-
[14]
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi- Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Augmented Generation. In Proceedings of the 2023 Conference on Empirical Meth- ods in Natural Language Processing . Association for...
2023
-
[15]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Ass...
2017
-
[16]
InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Asso...
2024
-
[17]
Jungo Kasai, Keisuke Sakaguchi, yoichi takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, and Ken- taro Inui. 2023. RealTime QA: What 's the Answer Right Now?. In Ad- vances in Neural Information Processing Systems . Curran Associates, I...
2023
-
[18]
Jaehyung Kim, Jaehyun Nam, Sangwoo Mo, Jongjin Park, Sang-Woo Lee, Min- joon Seo, Jung-Woo Ha, and Jinwoo Shin. 2024. SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs. In The Twelfth Interna- tional Conference on Learning Representations . https:...
2024
-
[19]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019
-
[20]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...
2020
-
[21]
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent Retrieval for Weakly Supervised Open Domain Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Optimizing Knowledge Integration in Retrieval-Augm...
2019
-
[22]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in ...
2020
-
[23]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12 (2024), 157–173
2024
-
[24]
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Ren Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. 2024. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. In Forty-first...
2024
-
[25]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023
-
[26]
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine- grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empiri...
2023
-
[27]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Ouyang Long, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021...
2021 arXiv
-
[28]
Yanming Liu, Xinyue Peng, Xuhong Zhang, Weihao Liu, Jianwei Yin, Jiannan Cao, and Tianyu Du. 2024. RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-Feedback. In Findings of the Associ- ation for Computational Linguistics: ACL 2024 . Asso...
2024
-
[29]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schul- man, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, a...
2022
-
[30]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimiza- tion: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 53728–...
2023
-
[31]
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don‘t Know: Unanswerable Questions for SQuAD. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics. https:...
2018
-
[32]
OpenAI. 2024. GPT-4 Technical Report. https://arxiv.org/abs/2303.08774
2024 arXiv
-
[33]
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-Context Retrieval-Augmented Lan- guage Models. Transactions of the Association for Computational Linguistics 11 (2023), 1316–1331. https://aclanthology.org/2023.tacl-1.75
2023
-
[34]
Nils Reimers and Iryna Gurevych. 2020. Making Monolingual Sentence Em- beddings Multilingual using Knowledge Distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics. https://arxiv.org/a...
2020 arXiv
-
[35]
Alireza Salemi, Surya Kallumadi, and Hamed Zamani. 2024. Optimization Meth- ods for Personalizing Large Language Models through Retrieval Augmentation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24...
2024
-
[36]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceed- ings of the 2016 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2383–2392
2016
-
[37]
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-Augmented Black-Box Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa...
2024
-
[38]
Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, and Soujanya Poria. 2024. Measuring and Enhancing Trustworthi- ness of LLMs in RAG through Grounded Attributions and Learning to Refuse. arXiv:2409.11242
2024 arXiv
-
[39]
Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020. Learning to summa- rize from human feedback. In Proceedings of the 34th International Conference on Neural Information Processing Systems (N...
2020
-
[40]
SciPhi-AI. 2024. SciPhi-Self-RAG-Mistral-7B-32k. https://huggingface.co/SciPhi/ SciPhi-Self-RAG-Mistral-7B-32k
2024
-
[41]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[42]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy...
2023 arXiv
-
[43]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[44]
Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V...
2024
-
[45]
Yile Wang, Peng Li, Maosong Sun, and Yang Liu. 2023. Self-Knowledge Guided Retrieval Augmentation for Large Language Models. In Findings of the Associa- tion for Computational Linguistics: EMNLP 2023 . Association for Computational Linguistics, 10303–10315. https://aclantholog...
2023
-
[46]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. InProceedings of the 36th International Conference on Neural Information Processin...
2022
-
[47]
Liu, and Matt Gardner
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing Multiple Choice Science Questions. In Proceedings of the 3rd Workshop on Noisy User- generated Text. Association for Computational Linguistics, 94–106
2017
-
[48]
In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge- Intensive Multi-Step Questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Toronto, Canada...
2023
-
[49]
Hongru Wang, Boyang Xue, Baohang Zhou, Tianhua Zhang, Cunxiang Wang, Guanhua Chen, Huimin Wang, and Kam fai Wong. 2024. Self-DC: When to retrieve and When to generate? Self Divide-and-Conquer for Compositional Unknown Questions
2024
-
[50]
Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua
-
[51]
Diji Yang, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo, Yawen Zhang, Jie Yang, and Yi Zhang. 2024. IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Info...
2024
-
[52]
Yichi Zhang, Zhuo Chen, Yin Fang, Yanxi Lu, Li Fangming, Wen Zhang, and Huajun Chen. 2024. Knowledgeable Preference Alignment for LLMs in Domain- specific Question Answering. In Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational ...
2024
-
[53]
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural Text Generation With Unlikelihood Training. In International Conference on Learning Representations . https://openreview.net/ forum?id=SJeYe0NtvH
2020
-
[54]
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. C-Pack: Packed Resources For General Chinese Embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24)...
2024
-
[56]
In Proceedings of the ACM Web Conference 2024 (WWW ’24)
Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive Tasks. In Proceedings of the ACM Web Conference 2024 (WWW ’24) . Association for Computing Machinery, 1362–1373. https://doi.org/10.1145/3589334.3645363
2024
-
[59]
Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, and Lu Wang. 2023. Merging Generated and Retrieved Knowledge for Open-Domain QA. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computa...
2023
-
[61]
Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. arXiv preprint arXiv:1909.08593 (2019). https: //arxiv.org/abs/1909.08593 Received 20 Februa...
2019 arXiv
-
[2020]
In Proceedings of the 37th International Conference on Machine Learning (ICML’20)
REALM: retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning (ICML’20) . JMLR.org, Article 368, 10 pages
-
[2023]
Atlas: few-shot learning with retrieval augmented language models. J. Mach. Learn. Res. 24, 1, Article 251 (Jan. 2023), 43 pages
2023
-
[2024]
In The Twelfth International Conference on Learning Representations
Self-RAG: Learning to Retrieve, Generate, and Critique through Self- Reflection. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=hSyW5go0v8
-
[4728]
https://aclanthology.org/2023.emnlp-main.286
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.