REVIEW 8 cited by
ChatGPT as a Factual Inconsistency Evaluator for Text Summarization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The performance of text summarization has been greatly boosted by pre-trained language models. A main concern of existing methods is that most generated summaries are not factually inconsistent with their source documents. To alleviate the problem, many efforts have focused on developing effective factuality evaluation metrics based on natural language inference, question answering, and syntactic dependency et al. However, these approaches are limited by either their high computational complexity or the uncertainty introduced by multi-component pipelines, resulting in only partial agreement with human judgement. Most recently, large language models(LLMs) have shown excellent performance in not only text generation but also language comprehension. In this paper, we particularly explore ChatGPT's ability to evaluate factual inconsistency under a zero-shot setting by examining it on both coarse-grained and fine-grained evaluation tasks including binary entailment inference, summary ranking, and consistency rating. Experimental results indicate that ChatGPT generally outperforms previous evaluation metrics across the three tasks, indicating its great potential for factual inconsistency evaluation. However, a closer inspection of ChatGPT's output reveals certain limitations including its preference for more lexically similar candidates, false reasoning, and inadequate understanding of instructions.
Forward citations
Cited by 8 Pith papers
-
Long-Form Information Alignment Evaluation Beyond Atomic Facts
A benchmark and evaluator for detecting misinformation composed entirely of truthful but reordered statements.
-
AdaptAgent: A Multi-agent, Domain-Guided Reasoning Framework for Code Adaptation
A multi-agent LLM pipeline that plans code adaptations using summarized intent, domain checklists, and sibling-method context outperforms single-shot prompting and repair baselines on Java adaptation examples.
-
Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions
Simplified prompting (Simple and ELI5) sharply reduces the number of word senses LLMs provide for homonyms, and DPO fine-tuning of Llama 3.1 8B restores much of that completeness.
-
Your Agent Can Defend Itself against Backdoor Attacks
A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.
-
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
A new benchmark of 23 LLM judges on 631 debate speeches shows large models approach human agreement but score lower, and LLM judges prefer GPT-4.1 speeches over human experts' without human verification.
-
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
REFLEX fine-tunes Phi-3-Mini on synthetic negative presentations to produce reference-free scores and actionable feedback for slide quality across coverage, redundancy, text-image alignment, and flow.
-
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
MADISSE assigns LLM evaluators random initial stances (faithful or unfaithful), has them debate in rounds, and adjudicates, improving summary faithfulness evaluation accuracy while introducing an annotated 'ambiguity'...
-
Prompt Engineering Guidelines for Using Large Language Models in Requirements Engineering
A literature review and three expert interviews yield a proposed mapping of prompt engineering guideline themes onto five requirements engineering activities, with no empirical validation of the mapping.
Discussion (0). Continue with ORCID to comment.