Pith. sign in

REVIEW 15 cited by

FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.13528 v2 pith:GF4JKYGN submitted 2023-07-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords factoolerrorsfactualgeneratedgenerativemodelschallengeschatgpt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of generative pre-trained models has facilitated the synthesis of high-quality text, but it has also posed challenges in identifying factual errors in the generated text. In particular: (1) A wider range of tasks now face an increasing risk of containing factual errors when handled by generative models. (2) Generated texts tend to be lengthy and lack a clearly defined granularity for individual facts. (3) There is a scarcity of explicit evidence available during the process of fact checking. With the above challenges in mind, in this paper, we propose FacTool, a task and domain agnostic framework for detecting factual errors of texts generated by large language models (e.g., ChatGPT). Experiments on four different tasks (knowledge-based QA, code generation, mathematical reasoning, and scientific literature review) show the efficacy of the proposed method. We release the code of FacTool associated with ChatGPT plugin interface at https://github.com/GAIR-NLP/factool .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Typed quantity verification exposes a canonical-equivalence blind spot in neural fact-checkers; Symbolic Augmentation fixes it (36.5%→98.2%) and transfers to SciFact-Open (+0.037 binary macro-F1).

  2. Before Agents Speak: Pre-hoc Failure Risk Inference in Multi-Agent Systems

    cs.CR 2026-07 conditional novelty 6.0 of 10

    HalluProp infers per-agent and system-level hallucination risk in multi-agent LLMs before interaction via role–query misalignment, topology-aware propagation, and differentiable Noisy-OR aggregation.

  3. Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.

  4. Reconsidering LLM Uncertainty Estimation Methods in the Wild

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.

  5. EMULATE: A Multi-Agent Framework for Determining the Veracity of Atomic Claims by Emulating Human Actions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EMULATE, a seven-agent LLM framework for fact-checking, reports top weighted-F1 on three benchmarks, beating FIRE by about 1 to 4 points.

  6. Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PhantomCircuit traces knowledge overshadowing to attention circuits during training and prunes circuit edges to recover the overshadowed answer.

  7. OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

    cs.AI 2025-08 conditional novelty 5.0 of 10

    OmniDPO extends direct preference optimization with audio-video alignment and modality-degradation preference pairs to reduce omni-modal hallucination.

  8. TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    TruthTorchLM is a new open-source library that standardizes 30+ LLM truthfulness prediction methods and benchmarks them on three datasets.

  9. RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new 6K-claim benchmark evaluates LLMs and multimodal LLMs on real-world fact-checking with an explicit 'unknown' option and shows web search and multimodal input improve performance.

  10. Inteligencia Artificial jur\'idica y el desaf\'io de la veracidad: an\'alisis de alucinaciones, optimizaci\'on de RAG y principios para una integraci\'on responsable

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Legal AI hallucination persists in commercial RAG tools (17-34%+ of queries), so the report argues the fix is consultative, source-citing system design plus mandatory human oversight, not better generative models.

  11. Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection

    cs.CL 2025-09 reject novelty 4.0 of 10

    Singular values of label-grouped n-gram frequency tensors are used as MLP features for hallucination detection, with reported gains on HaluEval that rely on label-aware grouping.

  12. FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Unanimous voting between NLI and chain-of-thought fact-checking yields scores closest to medical expert judgments on three of four tasks in the new FActBench benchmark.

  13. Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A multi-agent LLM pipeline with credibility-filtered full-text web retrieval reports better fact-checking F1 than four baselines on small benchmark subsamples.

  14. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

  15. AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions

    cs.AI 2025-09 conditional novelty 2.0 of 10

    A cross-domain vision paper that surveys AI-generated content and proposes research directions, without introducing new empirical results.

Pith tools