Pith. sign in

REVIEW 9 cited by

ChatGPT as a Factual Inconsistency Evaluator for Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.15621 v2 pith:QE7QABGT submitted 2023-03-27 cs.CL

classification cs.CL
keywords chatgptevaluationlanguagefactualinconsistencytexthoweverincluding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of text summarization has been greatly boosted by pre-trained language models. A main concern of existing methods is that most generated summaries are not factually inconsistent with their source documents. To alleviate the problem, many efforts have focused on developing effective factuality evaluation metrics based on natural language inference, question answering, and syntactic dependency et al. However, these approaches are limited by either their high computational complexity or the uncertainty introduced by multi-component pipelines, resulting in only partial agreement with human judgement. Most recently, large language models(LLMs) have shown excellent performance in not only text generation but also language comprehension. In this paper, we particularly explore ChatGPT's ability to evaluate factual inconsistency under a zero-shot setting by examining it on both coarse-grained and fine-grained evaluation tasks including binary entailment inference, summary ranking, and consistency rating. Experimental results indicate that ChatGPT generally outperforms previous evaluation metrics across the three tasks, indicating its great potential for factual inconsistency evaluation. However, a closer inspection of ChatGPT's output reveals certain limitations including its preference for more lexically similar candidates, false reasoning, and inadequate understanding of instructions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Long-Form Information Alignment Evaluation Beyond Atomic Facts

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A benchmark and evaluator for detecting misinformation composed entirely of truthful but reordered statements.

  2. AdaptAgent: A Multi-agent, Domain-Guided Reasoning Framework for Code Adaptation

    cs.SE 2026-08 conditional novelty 6.0 of 10

    A multi-agent LLM pipeline that plans code adaptations using summarized intent, domain checklists, and sibling-method context outperforms single-shot prompting and repair baselines on Java adaptation examples.

  3. Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Simplified prompting (Simple and ELI5) sharply reduces the number of word senses LLMs provide for homonyms, and DPO fine-tuning of Llama 3.1 8B restores much of that completeness.

  4. Your Agent Can Defend Itself against Backdoor Attacks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.

  5. Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark of 23 LLM judges on 631 debate speeches shows large models approach human agreement but score lower, and LLM judges prefer GPT-4.1 speeches over human experts' without human verification.

  6. Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback

    cs.CL 2025-05 conditional novelty 5.0 of 10

    REFLEX fine-tunes Phi-3-Mini on synthetic negative presentations to produce reference-free scores and actionable feedback for slide quality across coverage, redundancy, text-image alignment, and flow.

  7. Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

    cs.CL 2025-02 conditional novelty 5.0 of 10

    MADISSE assigns LLM evaluators random initial stances (faithful or unfaithful), has them debate in rounds, and adjudicates, improving summary faithfulness evaluation accuracy while introducing an annotated 'ambiguity'...

  8. GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing

    cs.CL 2025-02 conditional novelty 5.0 of 10

    GuideLLM is a modular system that steers autobiography interviews with a structured protocol, memory-graph questions, summarization, and emotion-aware responses, and it beats six LLM baselines in most automatic and hu...

  9. Prompt Engineering Guidelines for Using Large Language Models in Requirements Engineering

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A literature review and three expert interviews yield a proposed mapping of prompt engineering guideline themes onto five requirements engineering activities, with no empirical validation of the mapping.

Pith tools