REVIEW 4 cited by
AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Existing datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we introduce AVeriTeC, a new dataset of 4,568 real-world claims covering fact-checks by 50 different organizations. Each claim is annotated with question-answer pairs supported by evidence available online, as well as textual justifications explaining how the evidence combines to produce a verdict. Through a multi-round annotation process, we avoid common pitfalls including context dependence, evidence insufficiency, and temporal leakage, and reach a substantial inter-annotator agreement of $\kappa=0.619$ on verdicts. We develop a baseline as well as an evaluation scheme for verifying claims through several question-answering steps against the open web.
Forward citations
Cited by 4 Pith papers
-
MEDIAREF: A Public Knowledge Store for Media Background Checks
MEDIAREF is a public, updatable web-document store that lets LLMs generate media background checks more reproducibly and with higher fact recall than zero-shot generation alone.
-
FactIR: A Real-World Zero-shot Open-Domain Retrieval Benchmark for Fact-Checking
FactIR is a new real-world open-domain retrieval benchmark for fact-checking, used to show that lexical and sparse retrievers rival dense models while a clustering-trained dense retriever leads.
-
Towards Automated Fact-Checking of Real-World Claims: Exploring Task Formulation and Assessment with LLMs
In a benchmark of 17,856 PolitiFact claims, larger Llama-3 models and retrieved web evidence improve automated fact-checking accuracy and justification quality, though fine-grained labels remain difficult.
-
Evidence-Ledger Adjudication for Claim-Evidence Traceability
An evidence-ledger workflow labels claim-evidence pairs as supported/contradicted/missing/mixed and routes unsupported claims back to authors, reporting 0.676 accuracy over TF-IDF's 0.383 on a 2,335-row benchmark.
Discussion (0). Continue with ORCID to comment.