Pith. sign in

REVIEW 8 cited by

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.03200 v1 pith:T44GF7MR submitted 2025-01-06 cs.CL

classification cs.CL
keywords modelsleaderboardresponsesuserdocumentfactsgroundingjudge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce FACTS Grounding, an online leaderboard and associated benchmark that evaluates language models' ability to generate text that is factually accurate with respect to given context in the user prompt. In our benchmark, each prompt includes a user request and a full document, with a maximum length of 32k tokens, requiring long-form responses. The long-form responses are required to be fully grounded in the provided context document while fulfilling the user request. Models are evaluated using automated judge models in two phases: (1) responses are disqualified if they do not fulfill the user request; (2) they are judged as accurate if the response is fully grounded in the provided document. The automated judge models were comprehensively evaluated against a held-out test-set to pick the best prompt template, and the final factuality score is an aggregate of multiple judge models to mitigate evaluation bias. The FACTS Grounding leaderboard will be actively maintained over time, and contains both public and private splits to allow for external participation while guarding the integrity of the leaderboard. It can be found at https://www.kaggle.com/facts-leaderboard.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Hallucination Self-Play co-evolves a generator and detector from one base LLM via RLAIF and RLVR, lifting a 7B model to match advanced LLMs on RAGTruth faithfulness detection.

  2. DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs

    cs.CL 2025-06 conditional novelty 7.0 of 10

    The paper introduces a taxonomy and benchmark for knowledge conflicts in search-augmented LLMs, and experiments show that prompting for conflict type improves response quality.

  3. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    cs.AI 2026-02 reject novelty 6.0 of 10

    Nearly half of 60 widely used LLM benchmarks show saturation, which increases with age; private test access does not prevent saturation, while expert curation does.

  4. A Neurosymbolic Approach to Natural Language Formalization and Verification

    cs.CL 2025-11 conditional novelty 6.0 of 10

    A neurosymbolic guardrail reports 99.2% soundness on a 522-item policy QA benchmark, mainly by rejecting 84% of correct answers; unvetted real-world policies score 96.8%.

  5. StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A synthetic table-to-text pipeline that generates and validates key-value extraction benchmarks, revealing that LLM-generated reports keep numerical facts intact but are poorly machine-extractable.

  6. How Does Response Length Affect Long-Form Factuality

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Longer LLM responses have lower factual precision, and the main cause appears to be the model exhausting reliable knowledge on a topic, not error propagation or long context.

  7. AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Leading generative AIs are useful for introductory Islamic learning but unreliable as authorities on Fiqh, citations, and madhhab-sensitive rulings without human verification.

  8. Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

    cs.CL 2026-02 conditional novelty 5.0 of 10

    A step-level reinforcement-learning reward combining a process reward model with truncated resampling reduces chain-of-thought faithfulness hallucinations in small reasoning models.

Pith tools