Pith. sign in

REVIEW 1 cited by

Evaluation of Semantic Answer Similarity Metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.12664 v2 pith:E2PDNOTZ submitted 2022-06-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords similarityanswerevaluationmetricssemanticsystemsabilitybuild
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There are several issues with the existing general machine translation or natural language generation evaluation metrics, and question-answering (QA) systems are indifferent in that context. To build robust QA systems, we need the ability to have equivalently robust evaluation systems to verify whether model predictions to questions are similar to ground-truth annotations. The ability to compare similarity based on semantics as opposed to pure string overlap is important to compare models fairly and to indicate more realistic acceptance criteria in real-life applications. We build upon the first to our knowledge paper that uses transformer-based model metrics to assess semantic answer similarity and achieve higher correlations to human judgement in the case of no lexical overlap. We propose cross-encoder augmented bi-encoder and BERTScore models for semantic answer similarity, trained on a new dataset consisting of name pairs of US-American public figures. As far as we are concerned, we provide the first dataset of co-referent name string pairs along with their similarities, which can be used for training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Querying Large Automotive Software Models: Agentic vs. Direct LLM Approaches

    cs.SE 2025-06 conditional novelty 5.0 of 10

    A ReAct agent that reads a 13,572-line Ecore model via file tools matched direct full-context prompting on accuracy for the best models while using roughly 180 times fewer prompt tokens.

Pith tools