Pith. sign in

REVIEW 2 cited by

SIFiD: Reassess Summary Factual Inconsistency Detection with LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07557 v1 pith:6XOAXP6G submitted 2024-03-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords detectioninconsistencysummarylanguagellmsdocumentdocumentsfactual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensuring factual consistency between the summary and the original document is paramount in summarization tasks. Consequently, considerable effort has been dedicated to detecting inconsistencies. With the advent of Large Language Models (LLMs), recent studies have begun to leverage their advanced language understanding capabilities for inconsistency detection. However, early attempts have shown that LLMs underperform traditional models due to their limited ability to follow instructions and the absence of an effective detection methodology. In this study, we reassess summary inconsistency detection with LLMs, comparing the performances of GPT-3.5 and GPT-4. To advance research in LLM-based inconsistency detection, we propose SIFiD (Summary Inconsistency Detection with Filtered Document) that identify key sentences within documents by either employing natural language inference or measuring semantic similarity between summaries and documents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories

    cs.SE 2024-12 conditional novelty 7.0 of 10

    An empirical study of 1,262 prompt changes across 243 GitHub repositories shows that developers mainly add and modify prompt components during feature development, rarely document the changes, and sometimes introduce ...

  2. SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new benchmark built with executable phrase-level edits shows that most LLMs detect and explain factual inconsistencies in summaries only weakly, with the best model scoring 0.49 on the joint task.

Pith tools