Pith. sign in

REVIEW 1 cited by

How to Evaluate Your Dialogue Models: A Review of Approaches

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.01369 v1 pith:JJCAMTDR submitted 2021-08-03 cs.CL

classification cs.CL
keywords evaluationdialoguemethodmethodsanalysisapproachesautomaticbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating the quality of a dialogue system is an understudied problem. The recent evolution of evaluation method motivated this survey, in which an explicit and comprehensive analysis of the existing methods is sought. We are first to divide the evaluation methods into three classes, i.e., automatic evaluation, human-involved evaluation and user simulator based evaluation. Then, each class is covered with main features and the related evaluation metrics. The existence of benchmarks, suitable for the evaluation of dialogue techniques are also discussed in detail. Finally, some open issues are pointed out to bring the evaluation method into a new frontier.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems

    cs.CL 2025-01 conditional novelty 5.0 of 10

    The paper proposes a taxonomy of positive friction movements in dialogue and provides simulated and correlational evidence that they improve task success and user mental-state modeling.

Pith tools