Pith. sign in

REVIEW 13 cited by

Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07521 v3 pith:F4MVOAHG submitted 2023-10-11 cs.CL

classification cs.CL
keywords llmsfactualityfactualsurveychallengesdomainserrorsfacts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This survey addresses the crucial issue of factuality in Large Language Models (LLMs). As LLMs find applications across diverse domains, the reliability and accuracy of their outputs become vital. We define the Factuality Issue as the probability of LLMs to produce content inconsistent with established facts. We first delve into the implications of these inaccuracies, highlighting the potential consequences and challenges posed by factual errors in LLM outputs. Subsequently, we analyze the mechanisms through which LLMs store and process facts, seeking the primary causes of factual errors. Our discussion then transitions to methodologies for evaluating LLM factuality, emphasizing key metrics, benchmarks, and studies. We further explore strategies for enhancing LLM factuality, including approaches tailored for specific domains. We focus two primary LLM configurations standalone LLMs and Retrieval-Augmented LLMs that utilizes external data, we detail their unique challenges and potential enhancements. Our survey offers a structured guide for researchers aiming to fortify the factual reliability of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conversational AI increases political knowledge as effectively as self-directed internet search

    cs.HC 2025-09 reject novelty 7.0 of 10

    Using survey and experimental data, the paper reports that conversational AI is a common source of political information in the UK and that its effect on belief in true versus false statements statistically does not e...

  2. RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A mixture-of-experts LLM trained with reinforcement learning to perform retrieval from its own parametric memory can replace external retrieval in some settings, at lower latency.

  3. An Automated Attack Investigation Approach Leveraging Threat-Knowledge-Augmented Large Language Models

    cs.CR 2025-09 reject novelty 6.0 of 10

    ANANKE reports 97.1% TPR and 0.2% FPR for LLM-based attack investigation, but its central evaluation is compromised by knowledge-base overlap with test scenarios.

  4. RewardAnything: Generalizable Principle-Following Reward Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RewardAnything follows natural-language reward principles at inference time and, with the new RABench benchmark, demonstrates that principle-conditioned listwise training beats fixed-preference reward models on held-o...

  5. How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new GraphRAG evaluation framework using graph-grounded questions and bias-correction yields much smaller win rates than earlier reports, casting doubt on reported GraphRAG gains.

  6. Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A controlled study shows RETRO-style models need a minimum input-neighbor overlap to activate, and paraphrased synthetic context can speed up training by about 40% at a small perplexity cost.

  7. SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence

    cs.CL 2025-02 conditional novelty 6.0 of 10

    SelfElicit uses deep-layer attention to automatically highlight relevant evidence sentences in the input context, yielding consistent QA accuracy gains across six instruction-tuned LLMs.

  8. Leveraging Knowledge Graphs and LLMs for Structured Generation of Misinformation

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A knowledge graph-guided LLM pipeline generates misinformation by replacing objects in true triplets with structurally similar ones, and LLM judges often fail to flag the fakes as fake.

  9. GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A modular zero-shot KB-VQA framework using Grounding DINO, dual captioners, semantic caption filtering, and LLM prompting reports new state-of-the-art numbers on OK-VQA, A-OKVQA, and VQAv2.

  10. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

    cs.HC 2025-02 conditional novelty 5.0 of 10

    Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.

  11. Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Pre-fine-tuning scores on a three-task diagnostic can predict the direction of post-fine-tuning change in small LLMs for cybersecurity QA, but not the magnitude or rank-preservation, which is regime-dependent.

  12. Confidence Estimation for Text-to-SQL in Large Language Models

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Consistency-based methods are the most reliable confidence signal for text-to-SQL in black-box LLMs, and executing queries against a database adds a useful correctness signal.

  13. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

Pith tools