Pith. sign in

REVIEW 2 cited by

PEDANTS: Cheap but Effective and Interpretable Answer Equivalence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11161 v5 pith:XEGI62BJ submitted 2024-02-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords answercurrentdatasetsevaluationexpensiveinterpretablellmsonly
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Question answering (QA) can only make progress if we know if an answer is correct, but current answer correctness (AC) metrics struggle with verbose, free-form answers from large language models (LLMs). There are two challenges with current short-form QA evaluations: a lack of diverse styles of evaluation data and an over-reliance on expensive and slow LLMs. LLM-based scorers correlate better with humans, but this expensive task has only been tested on limited QA datasets. We rectify these issues by providing rubrics and datasets for evaluating machine QA adopted from the Trivia community. We also propose an efficient, and interpretable QA evaluation that is more stable than an exact match and neural methods(BERTScore).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MinosEval improves open-ended QA evaluation by sorting questions into factoid and non-factoid and applying tailored scoring, outperforming baselines on four datasets.

  2. Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    PrefBERT, a 150M-parameter reward model trained on human quality ratings, outperforms ROUGE-L and BERTScore as a GRPO reward signal for open-ended long-form generation.

Pith tools