Pith. sign in

REVIEW 9 cited by

FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03214 v2 pith:UNO3AZK5 submitted 2023-10-05 cs.CL

classification cs.CL
keywords freshqamodelsquestionsanswersfreshpromptknowledgesearchworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most large language models (LLMs) are trained once and never updated; thus, they lack the ability to dynamically adapt to our ever-changing world. In this work, we perform a detailed study of the factuality of LLM-generated text in the context of answering questions that test current world knowledge. Specifically, we introduce FreshQA, a novel dynamic QA benchmark encompassing a diverse range of question and answer types, including questions that require fast-changing world knowledge as well as questions with false premises that need to be debunked. We benchmark a diverse array of both closed and open-source LLMs under a two-mode evaluation procedure that allows us to measure both correctness and hallucination. Through human evaluations involving more than 50K judgments, we shed light on limitations of these models and demonstrate significant room for improvement: for instance, all models (regardless of model size) struggle on questions that involve fast-changing knowledge and false premises. Motivated by these results, we present FreshPrompt, a simple few-shot prompting method that substantially boosts the performance of an LLM on FreshQA by incorporating relevant and up-to-date information retrieved from a search engine into the prompt. Our experiments show that FreshPrompt outperforms both competing search engine-augmented prompting methods such as Self-Ask (Press et al., 2022) as well as commercial systems such as Perplexity.AI. Further analysis of FreshPrompt reveals that both the number of retrieved evidences and their order play a key role in influencing the correctness of LLM-generated answers. Additionally, instructing the LLM to generate concise and direct answers helps reduce hallucination compared to encouraging more verbose answers. To facilitate future work, we release FreshQA at github.com/freshllms/freshqa and commit to updating it at regular intervals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

    cs.CL 2026-08 accept novelty 6.0 of 10

    Spatial-memory staleness is a measurable safety failure for VLM agents: stale memory increases deaths, and visual auditing of stale entries is highly model-dependent.

  2. Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    DOVE measures LLM cultural value alignment via a rate-distortion value codebook and unbalanced optimal transport between human and model open-ended text distributions.

  3. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  4. Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An automated LLM-based evaluator finds that eight LLMs rarely hallucinate factors in legal argument generation but often omit relevant factors and usually fail to abstain when no common ground exists.

  5. MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A pre-training task called RAMP, where models practice searching to fill masked text spans, improves downstream agentic open-domain QA performance across Qwen and LLaMA models.

  6. InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    InfoDeepSeek is a 245-question benchmark that measures how well AI agents seek information on the live web, with new metrics for answer accuracy, evidence quality, and compactness.

  7. MedBrowseComp: Benchmarking Medical Deep Research and Computer Use

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark of more than 1,000 multi-hop medical browsing questions shows that even the best deep-research and computer-use AI agents answer fewer than half correctly.

  8. A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models

    cs.IR 2025-07 reject novelty 3.0 of 10

    A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.

  9. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools