Pith. sign in

REVIEW 2 cited by

USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00456 v1 pith:JNV762FF submitted 2020-05-01 cs.CL cs.LG

classification cs.CLcs.LG
keywords dialogevaluationmetricunsuperviseddesirablegenerationmetricsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog. USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog. USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, system-level: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0). USR additionally produces interpretable measures for several desirable properties of dialog.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    HDCEval splits medical answer grading into relevance, correctness, and expression checks, uses reward-token-trained expert models, and reports improved agreement with human doctors.

  2. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

Pith tools