REVIEW 7 cited by
Rethinking with Retrieval: Faithful Large Language Model Inference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite the success of large language models (LLMs) in various natural language processing (NLP) tasks, the stored knowledge in these models may inevitably be incomplete, out-of-date, or incorrect. This motivates the need to utilize external knowledge to assist LLMs. Unfortunately, current methods for incorporating external knowledge often require additional training or fine-tuning, which can be costly and may not be feasible for LLMs. To address this issue, we propose a novel post-processing approach, rethinking with retrieval (RR), which retrieves relevant external knowledge based on the decomposed reasoning steps obtained from the chain-of-thought (CoT) prompting. This lightweight approach does not require additional training or fine-tuning and is not limited by the input length of LLMs. We evaluate the effectiveness of RR through extensive experiments with GPT-3 on three complex reasoning tasks: commonsense reasoning, temporal reasoning, and tabular reasoning. Our results show that RR can produce more faithful explanations and improve the performance of LLMs.
Forward citations
Cited by 7 Pith papers
-
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law
SafeWork-R1 shows that a staged RL pipeline with safety, value, and knowledge verifiers can improve both safety and general reasoning scores over a base multimodal model.
-
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.
-
GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
GainRAG aligns retriever and LLM preferences by training a selector on contrastive-perplexity 'gain' signals plus a pseudo-passage fallback, improving RAG accuracy on six QA datasets.
-
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
InfoDeepSeek is a 245-question benchmark that measures how well AI agents seek information on the live web, with new metrics for answer accuracy, evidence quality, and compactness.
-
Enhancing Health Information Retrieval with RAG by Prioritizing Topical Relevance and Factual Accuracy
A three-stage RAG pipeline generates a cited summary (GenText) from PubMed Central passages and ranks health documents by topical relevance plus alignment with that summary, outperforming baselines on CLEF eHealth and...
-
Deploying AI for Signal Processing education: Selected challenges and intriguing opportunities
AI can be used to generate interactive signal processing courseware, but the paper offers no evidence that students learn better from it.
-
TeroSeek: An AI-Powered Knowledge Base and Retrieval Generation Platform for Terpenoid Research
TeroSeek couples a curated terpenoid literature knowledge base with retrieval-augmented generation to answer domain questions with citations, reporting higher accuracy than general-purpose LLMs on a 41-question test set.
Discussion (0). Continue with ORCID to comment.