Pith. sign in

REVIEW 12 cited by

CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.00614 v2 pith:KPLRRMHG submitted 2020-12-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords claimsclimatedatasetclimate-fevercommunityfeverlanguagereal-world
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce CLIMATE-FEVER, a new publicly available dataset for verification of climate change-related claims. By providing a dataset for the research community, we aim to facilitate and encourage work on improving algorithms for retrieving evidential support for climate-specific claims, addressing the underlying language understanding challenges, and ultimately help alleviate the impact of misinformation on climate change. We adapt the methodology of FEVER [1], the largest dataset of artificially designed claims, to real-life claims collected from the Internet. While during this process, we could rely on the expertise of renowned climate scientists, it turned out to be no easy task. We discuss the surprising, subtle complexity of modeling real-world climate-related claims within the \textsc{fever} framework, which we believe provides a valuable challenge for general natural language understanding. We hope that our work will mark the beginning of a new exciting long-term joint effort by the climate science and AI community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mutual Linearity in and out of Stationarity for Markov Jump Processes: A Trajectory-Based Approach

    cond-mat.stat-mech 2026-04 unverdicted novelty 7.0 of 10

    Trajectory-level linear response yields mutual linearity of observables under single-edge rate perturbation for Markov jump processes, including non-stationary state and counting observables.

  2. Tailored untruths: How personalisation challenges LLM safeguards

    cs.CL 2025-10 conditional novelty 7.0 of 10

    A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.

  3. ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A large-scale benchmark shows that leading multimodal language models still underperform expert humans at verifying climate claims from scientific charts.

  4. Climate Finance Bench

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Climate Finance Bench releases 330 expert-validated QA pairs on 33 climate reports and shows that retrieval quality, not model capacity, is the main accuracy bottleneck.

  5. Score-Only Distillation for Compact Dense Retrieval

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Score-only distillation with a row-centered all-pairs PairMSE objective lets 0.6B bi-encoders recover up to 50% of the base-to-teacher retrieval gap under matched protocols.

  6. Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

    cs.IR 2025-05 conditional novelty 5.0 of 10

    Contrastive fine-tuning often degrades strong dense retrievers, while combining cross-encoder listwise distillation with diverse synthetic queries consistently improves them.

  7. Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ClimateEval unifies 25 climate-related NLP tasks into one benchmark and shows that open-source LLMs gain from few-shot examples but lag on misinformation and fine-grained entity recognition.

  8. Evidence-Ledger Adjudication for Claim-Evidence Traceability

    cs.AI 2026-07 conditional novelty 4.0 of 10

    An evidence-ledger workflow labels claim-evidence pairs as supported/contradicted/missing/mixed and routes unsupported claims back to authors, reporting 0.676 accuracy over TF-IDF's 0.383 on a 2,335-row benchmark.

  9. HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers

    cs.IR 2025-09 conditional novelty 4.0 of 10

    By first fusing multiple retrievers within labeled and unlabeled sources with RRF, then merging z-score normalized lists, HF-RAG improves fact-verification F1 in-domain and out-of-domain.

  10. Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A partial reproduction of Vec2Text confirms that embedding inversion is real, and shows that 8-bit quantization of stored embeddings is a lightweight defense that preserves retrieval while reducing reconstruction.

  11. Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems

    cs.IR 2025-05 conditional novelty 4.0 of 10

    A reranker fine-tuned on hard negatives selected by two cosine-distance criteria outperforms older negative sampling methods on enterprise and domain-specific retrieval benchmarks.

  12. Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models

    cs.CL 2025-05 reject novelty 3.0 of 10

    Small language models fine-tuned on GPT-4 causal explanations score high on a new teacher-similarity metric, but the paper provides no independent evidence that causal reasoning was transferred.

Pith tools