REVIEW 8 cited by
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce FACTS Grounding, an online leaderboard and associated benchmark that evaluates language models' ability to generate text that is factually accurate with respect to given context in the user prompt. In our benchmark, each prompt includes a user request and a full document, with a maximum length of 32k tokens, requiring long-form responses. The long-form responses are required to be fully grounded in the provided context document while fulfilling the user request. Models are evaluated using automated judge models in two phases: (1) responses are disqualified if they do not fulfill the user request; (2) they are judged as accurate if the response is fully grounded in the provided document. The automated judge models were comprehensively evaluated against a held-out test-set to pick the best prompt template, and the final factuality score is an aggregate of multiple judge models to mitigate evaluation bias. The FACTS Grounding leaderboard will be actively maintained over time, and contains both public and private splits to allow for external participation while guarding the integrity of the leaderboard. It can be found at https://www.kaggle.com/facts-leaderboard.
Forward citations
Cited by 8 Pith papers
-
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator
Hallucination Self-Play co-evolves a generator and detector from one base LLM via RLAIF and RLVR, lifting a 7B model to match advanced LLMs on RAGTruth faithfulness detection.
-
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
The paper introduces a taxonomy and benchmark for knowledge conflicts in search-augmented LLMs, and experiments show that prompting for conflict type improves response quality.
-
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Nearly half of 60 widely used LLM benchmarks show saturation, which increases with age; private test access does not prevent saturation, while expert curation does.
-
A Neurosymbolic Approach to Natural Language Formalization and Verification
A neurosymbolic guardrail reports 99.2% soundness on a 522-item policy QA benchmark, mainly by rejecting 84% of correct answers; unvetted real-world policies score 96.8%.
-
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
A synthetic table-to-text pipeline that generates and validates key-value extraction benchmarks, revealing that LLM-generated reports keep numerical facts intact but are poorly machine-extractable.
-
How Does Response Length Affect Long-Form Factuality
Longer LLM responses have lower factual precision, and the main cause appears to be the model exhausting reliable knowledge on a topic, not error propagation or long context.
-
AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge
Leading generative AIs are useful for introductory Islamic learning but unreliable as authorities on Fiqh, citations, and madhhab-sensitive rulings without human verification.
-
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models
A step-level reinforcement-learning reward combining a process reward model with truncated resampling reduces chain-of-thought faithfulness hallucinations in small reasoning models.
Discussion (0). Sign in to comment.