Pith. sign in

REVIEW 2 cited by

FLEEK: Factual Error Detection and Correction with Evidence Retrieved from External Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17119 v1 pith:Z6FWMK5K submitted 2023-10-26 cs.CL

classification cs.CL
keywords factualerrorsfleekdetectionevidenceexternalknowledgeclaims
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Detecting factual errors in textual information, whether generated by large language models (LLM) or curated by humans, is crucial for making informed decisions. LLMs' inability to attribute their claims to external knowledge and their tendency to hallucinate makes it difficult to rely on their responses. Humans, too, are prone to factual errors in their writing. Since manual detection and correction of factual errors is labor-intensive, developing an automatic approach can greatly reduce human effort. We present FLEEK, a prototype tool that automatically extracts factual claims from text, gathers evidence from external knowledge sources, evaluates the factuality of each claim, and suggests revisions for identified errors using the collected evidence. Initial empirical evaluation on fact error detection (77-85\% F1) shows the potential of FLEEK. A video demo of FLEEK can be found at https://youtu.be/NapJFUlkPdQ.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Aggregate Text2SQL benchmark numbers are distorted by ambiguous single labels and by the SQL-equivalence match functions, a problem the paper organizes into a taxonomy with concrete Spider examples.

  2. A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models

    cs.IR 2025-07 reject novelty 3.0 of 10

    A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.

Pith tools