Pith. sign in

REVIEW 1 cited by

A Closer Look at Claim Decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11903 v1 pith:SVUJ7YBI submitted 2024-03-18 cs.CL

classification cs.CL
keywords decompositiontextmethodsapproachclaimfactscoregeneratedllm-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As generated text becomes more commonplace, it is increasingly important to evaluate how well-supported such text is by external knowledge sources. Many approaches for evaluating textual support rely on some method for decomposing text into its individual subclaims which are scored against a trusted reference. We investigate how various methods of claim decomposition -- especially LLM-based methods -- affect the result of an evaluation approach such as the recently proposed FActScore, finding that it is sensitive to the decomposition method used. This sensitivity arises because such metrics attribute overall textual support to the model that generated the text even though error can also come from the metric's decomposition step. To measure decomposition quality, we introduce an adaptation of FActScore, which we call DecompScore. We then propose an LLM-based approach to generating decompositions inspired by Bertrand Russell's theory of logical atomism and neo-Davidsonian semantics and demonstrate its improved decomposition quality over previous methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Large language models are less logically consistent when hypotheses are decomposed into atomic sub-problems, and a new inferential-consistency metric quantifies how consistently models handle the same fact in differen...

Pith tools