Pith. sign in

REVIEW 4 cited by

VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19276 v1 pith:NSENTBHS submitted 2024-06-27 cs.CL

classification cs.CL
keywords veriscorelong-formtasksgenerationacrossclaimsdifferentfactuality
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing metrics for evaluating the factuality of long-form text, such as FACTSCORE (Min et al., 2023) and SAFE (Wei et al., 2024), decompose an input text into "atomic claims" and verify each against a knowledge base like Wikipedia. These metrics are not suitable for most generation tasks because they assume that every claim is verifiable (i.e., can plausibly be proven true or false). We address this issue with VERISCORE, a metric for diverse long-form generation tasks that contain both verifiable and unverifiable content. VERISCORE can be effectively implemented with either closed or fine-tuned open-weight language models, and human evaluation confirms that VERISCORE's extracted claims are more sensible than those from competing methods across eight different long-form tasks. We use VERISCORE to evaluate generations from 16 different models across multiple long-form tasks and find that while GPT-4o is the best-performing model overall, open-weight models such as Mixtral-8x22 are closing the gap. We show that an LM's VERISCORE on one task (e.g., biography generation) does not necessarily correlate to its VERISCORE on a different task (e.g., long-form QA), highlighting the need for expanding factuality evaluation across tasks with varying fact density.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

    cs.CL 2025-01 conditional novelty 6.0 of 10

    FACTS Grounding is a benchmark and leaderboard that scores LLMs on producing long-form answers fully grounded in up to 32k-token documents, using a validated panel of judge models.

  2. Towards Long Context Hallucination Detection

    cs.CL 2025-04 conditional novelty 5.0 of 10

    A decomposition-and-aggregation encoder architecture, trained on a new GPT-4o-injected BookSum dataset, outperforms LLM baselines on long-context hallucination detection while running much faster.

  3. Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A dynamic RAG pipeline generates situation awareness reports for peacebuilding from GDELT, ACLED, ReliefWeb, and World Bank data, evaluated by NLP metrics, UNDP experts, and LLM judges.

  4. Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey

    cs.CL 2025-04 conditional novelty 3.0 of 10

    A review that organizes RAG evaluation into internal and external categories, catalogs dozens of benchmarks, and analyzes evaluation practices in 582 conference papers.

Pith tools