REVIEW 2 cited by
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Ensuring factual consistency between the summary and the original document is paramount in summarization tasks. Consequently, considerable effort has been dedicated to detecting inconsistencies. With the advent of Large Language Models (LLMs), recent studies have begun to leverage their advanced language understanding capabilities for inconsistency detection. However, early attempts have shown that LLMs underperform traditional models due to their limited ability to follow instructions and the absence of an effective detection methodology. In this study, we reassess summary inconsistency detection with LLMs, comparing the performances of GPT-3.5 and GPT-4. To advance research in LLM-based inconsistency detection, we propose SIFiD (Summary Inconsistency Detection with Filtered Document) that identify key sentences within documents by either employing natural language inference or measuring semantic similarity between summaries and documents.
Forward citations
Cited by 2 Pith papers
-
Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories
An empirical study of 1,262 prompt changes across 243 GitHub repositories shows that developers mainly add and modify prompt components during feature development, rarely document the changes, and sometimes introduce ...
-
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
A new benchmark built with executable phrase-level edits shows that most LLMs detect and explain factual inconsistencies in summaries only weakly, with the best model scoring 0.49 on the joint task.
Discussion (0). Continue with ORCID to comment.