Pith. sign in

REVIEW 3 cited by

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11972 v3 pith:JQNSD25V submitted 2025-04-16 cs.CL

classification cs.CL
keywords llm-as-a-judgedatasetsextractivemodeloftenpromptacrossbias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using large language models (LLMs) as judges (LLM-as-a-judge), yet they often lack comprehensive evaluation across datasets and overlook key factors such as sensitivity to answer types, prompt variations, and self-preference bias. In this work, we conduct a systematic study of LLM-as-a-judge across four extractive QA datasets and various prompt variations, assessing multiple LLM families in both answering and judging roles. Our results show that LLM-as-a-judge judgments correlate much more strongly with human evaluations than EM (0.22) and F1 (0.40), achieving correlations up to 0.85 with open-source models. Further analysis reveals that LLM-as-a-judge performs particularly well on number-related answers but faces challenges with more complex types, such as job titles. Contrary to findings in other NLP tasks, we observe no self-preference bias, even when the same model serves as both QA model and judge. Finally, we find that prompt phrasing has minimal impact, and zero-shot, context-free judging often yields the best evaluation performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A systematic human audit of SocialIQa, FauxPas-EAI and ToMi shows that benchmark scores are inflated or distorted by data flaws, rigid scoring, and sensitivity to phrasing.

  2. Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    SA-RAG uses tables as a compact intermediate representation for RAG, introduces a quality-aware table metadata generation framework, and reports outperformance over baselines on two noisy real-world datasets.

  3. Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Coordinated misleading text descriptions of video, audio, and meaning flip the appropriateness labels assigned by most multimodal LLMs in about 90% of test videos.

Pith tools