REVIEW 1 cited by
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation. In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer. Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST. We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages. We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect.
Forward citations
Cited by 1 Pith paper
-
Absolute Evaluation Measures for Machine Learning: A Survey
A survey compiles bounded absolute evaluation metrics for classification, clustering, and ranking and proposes decision trees for metric selection, but several formulas are reproduced incorrectly.
Discussion (0). Continue with ORCID to comment.