Pith. sign in

REVIEW 3 major objections 4 minor 81 references

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Human annotators are not consistently better than automatic metrics in machine-translation evaluation: state-of-the-art metrics often rank on par with or above human baselines.

desk verdict A careful empirical study that puts human baselines from multiple protocols into WMT meta-evaluation, with honest limitations; the 'parity' result is suggestive but not yet established. read the letter →

arxiv 2506.19571 v1 pith:HSBNYZIZ submitted 2025-06-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords machinetranslationevaluationmeta-evaluationhumanparitybaselinesinter-annotatoragreementsoftpairwiseaccuracytiecalibrationWMTsharedtask
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether automatic machine-translation metrics have reached human parity as evaluators, and returns a qualified yes. When human annotators are entered into the same leaderboards used to rank metrics, state-of-the-art metrics often match or exceed human baselines: human evaluators share the top statistical cluster with metrics under system-level Soft Pairwise Accuracy and are frequently surpassed under segment-level tie-calibrated Pairwise Accuracy ($acc^*_{eq}$). The comparison draws on seven WMT test sets with overlapping annotations from several human protocols, and raters are partitioned into disjoint evaluator groups so that human–human agreement is not inflated. The authors then argue that these results are not enough to declare human parity, and that they make further progress in MT evaluation hard to measure.

What carries the argument

The machinery is a human-baseline meta-evaluation setup. Each test set is restricted to the largest subset of segments that can be covered by evaluators built from disjoint sets of raters, found by solving an integer linear program, so that no rater contributes to both the ground truth and a baseline. The resulting human evaluators are then ranked against automatic metrics with Soft Pairwise Accuracy (SPA), which rewards an evaluator for expressing confidence levels close to the ground truth when ranking MT systems, and with tie-calibrated Pairwise Accuracy ($acc^*_{eq}$), which counts segment-level pairwise agreements while calibrating ties per evaluator. The interplay of these two measures carries the argument: human evaluators look near the top under SPA but drop under $acc^*_{eq}$, and the paper attributes the drop to human scores being discrete rather than continuous.

What would settle it

Take a test set of translations with injected, expert-verified errors that human raters reliably catch—gender agreement, number inflection, named entities, word-sense ambiguities—and recompute SPA and $acc^*_{eq}$ with human baselines and top metrics; if human evaluators then consistently rank above every metric, the reported parity was an artifact of easy test sets and discrete human scales.

Watch

Extended reading notes

Core claim

The central claim is that, in current MT meta-evaluation, the human reference no longer separates people from machines. Treating one MQM-based human evaluator as ground truth and other human protocols as baselines, the authors rank every evaluator—human and automatic—on the same WMT 2024 meta-evaluation measures. Human baselines do not consistently rank higher than automatic metrics; under SPA they often sit in the same significance cluster as top metrics, and under $acc^*_{eq}$ they are frequently outranked. The finding is robust to swapping the ground-truth evaluator, but the authors caution that it does not establish equivalence: a metric that ranks above a human may simply align more closely with the score distribution of the protocol used as gold, and current test sets may be too easy for the gap to be meaningful.

Load-bearing premise

The load-bearing premise is that agreement with one chosen MQM-based human evaluator is a valid yardstick for ranking both automatic metrics and human raters who used other protocols; if that premise fails, metrics that outrank humans may just be matching the MQM score distribution rather than evaluating better.

Editorial extensions

If this is right

  • A metric that ranks above a human baseline in a WMT-style leaderboard no longer demonstrates that it evaluates better than a person; it may only show closer agreement with the MQM protocol used as gold.
  • Rankings under $acc^*_{eq}$ systematically penalize human raters for producing discrete scores, so cross-protocol comparisons of human and automatic evaluators are confounded by score granularity.
  • If current test sets are as easy as the fluency-only sentinel metric result suggests, then measured parity may vanish once metrics are tested on adversarial or out-of-domain translations.
  • The field's ability to track improvement in MT evaluation is at risk: once top metrics sit at the human ceiling, a higher ranking is ambiguous and no longer a clear sign of progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the human reference is saturated, optimizing a metric for the leaderboard becomes equivalent to optimizing it for the specific human raters behind the gold annotations, which rewards fitting the benchmark rather than evaluation quality.
  • A sharper test of parity would inject expert-verified errors (wrong gender, wrong number, named-entity errors, word-sense ambiguities) into translations and check whether human baselines still lose to metrics on those segments; the paper's data cannot answer this.
  • Building a consensus ground truth from several protocols instead of a single MQM evaluator, or letting human raters give continuous severity scores, might move human baselines to the other side of the parity line.
  • The gap between SPA and $acc^*_{eq}$ suggests the human disadvantage is largely a segment-level, scale-granularity phenomenon; a human protocol with finer-grained scores could plausibly reverse the ranking.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper re-runs the WMT 2024 meta-evaluation protocols on seven test sets from the 2020, 2022, 2023, and 2024 WMT Metrics Shared Tasks, adding human annotators as evaluators alongside automatic metrics. Using a single MQM-based human evaluator as ground truth, the authors compute Soft Pairwise Accuracy (SPA) and Pairwise Accuracy with Tie Calibration (acc*_eq) for both metrics and human baselines from other protocols (ESA, pSQM, DA+SQM). The central claim is that automatic metrics often rank in the same statistical significance cluster as human baselines under SPA and frequently surpass them under acc*_eq, which the paper interprets as suggesting, while explicitly cautioning against, human parity in MT evaluation. The paper also discusses the limits of measuring progress when the human reference no longer separates evaluator types.

Significance. If the empirical result is robust, it has important consequences: current WMT meta-evaluation may no longer be able to distinguish automatic metrics from human raters, and incremental metric improvements become ambiguous. The paper makes a useful contribution by bringing human baselines into the standard meta-evaluation framework and by explicitly raising the interpretation problem. Strengths include the use of official WMT meta-evaluation measures, statistical significance clusters on rankings, and a robustness appendix that varies the ground-truth evaluator; the authors also release their code. However, the central claim is currently conditional on the choice of a single gold evaluator and on the treatment of discrete human scores under acc*_eq, which the paper itself identifies but does not resolve experimentally.

major comments (3)
  1. [§2.2 and Appendix F] The ground-truth evaluator is always a single human evaluator (MQM in the main results; pSQM-1, DA+SQM, MQM-2023-2/3, or ESA in Appendix F), never a consensus of multiple raters, and the 2024 EN→ES test set has exactly one MQM and one ESA evaluator. As the authors themselves ask in §4.1, a high ranking may simply reflect closer alignment with the score distribution of the chosen gold protocol or rater rather than better evaluation capability; this is exactly the sort of protocol-mismatch artifact that could put human baselines below metrics. To support the headline claim that human baselines are not consistently superior, the paper should add a consensus-based gold (e.g., averaging the multiple MQM evaluators available for 2020, 2022, and 2023) or a leave-one-out analysis across gold evaluators, and show that the relative ordering of metrics and human baselines is stable under those conditions.
  2. [§4, Appendix C.2, Table 2] The acc*_eq measure systematically ranks human evaluators far lower than SPA does (e.g., pSQM-2 at rank 9 in 2020 EN→DE, DA+SQM at rank 16 in 2022 EN→DE, DA+SQM at rank 14 in 2023 EN→DE), and this is the main source of the 'metrics surpass humans' evidence. The authors attribute this to Perrella et al. (2024b)'s finding that tie calibration favors continuous-scale evaluators while human protocols produce discrete scores, but they do not quantify or control for this scale mismatch. A control experiment that rounds metric scores to the integer granularity of human protocols, or that applies an identical tie-calibration policy to all evaluators, would show whether the acc*_eq gaps are an artifact of score granularity rather than of evaluation quality; without this control, the acc*_eq results do not by themselves establish that metrics outperform human baselines.
  3. [Table 1, footnote 3, Table 6] The 2023 EN→DE test set contains only 145 segments after filtering, yet it is one of the two test sets where human baselines fall most consistently below automatic metrics under both SPA and acc*_eq. The authors acknowledge this small sample in footnote 3, but the main text does not report whether the Table 6 rankings are stable when the test set is enlarged to the 376 segments used in Appendix F (Tables 9–12, with different gold evaluators). Given that the 2023 EN→DE result is a key piece of evidence for the paper's central claim, the paper should either present the Table 6 rankings on the larger segment set with the same MQM gold, or provide confidence intervals for the rank differences, to demonstrate that the 145-segment fragility does not drive the headline result.
minor comments (4)
  1. [§2.2] The sentence 'Following standard practice in the literature ... we designate evaluators derived from the MQM annotations ... as the ground truth' correctly notes the convention, but the paper should also acknowledge that in 2020 the pSQM protocol was also professional and that the choice of MQM as the gold is a substantive modeling decision rather than purely a default.
  2. [Footnote 3] The caution about the 145-segment 2023 EN→DE set appears only in a footnote at the end of the Discussion; it should be prominently placed near Table 1 or Table 6, since it directly bears on the main result.
  3. [Table 2 and Appendix E] The use of 'Acc.' and 'Rank' in the captions is clear, but the superscripting of acc*_eq is inconsistent in several places; please unify the notation.
  4. [Section 3] The phrase 'human evaluators do not consistently rank higher than automatic metrics' could be made more precise by adding 'in the specific test sets and meta-evaluation measures considered here', to avoid overgeneralization.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the human-parity measurement is a computed meta-evaluation against an explicit human gold, with protocol-mismatch caveats stated openly; no fitted parameter or load-bearing self-citation is present.

full rationale

The paper does not fit parameters to the target claim and then report them as predictions. It takes published WMT human annotations and published metrics, applies the WMT 2024 meta-evaluation measures SPA and acc*_eq, and ranks all evaluators against one designated human ground truth. That gold choice is explicit in Section 2.2 ('we designate one human evaluator as ground truth while the others serve as human baselines') and is varied in Appendix F, so the finding that metrics often rank at or above human baselines is a computed outcome, not an identity. The concern that a higher metric ranking may mean 'merely align[ing] more closely with the score distribution of the MQM protocol' is raised by the authors themselves in Section 4.1 as a reason for caution; the paper does not conceal the yardstick problem but flags it as a limitation. The mentions of Perrella et al. (2024a,b) are contextual: one credits prior joint evaluation of metrics and humans, and the other is invoked to explain the acc*_eq discrepancy, but neither is used as an unverified premise that forces the result. No equation reduces to its own input, and no parameter fitted to a subset of data is renamed as a prediction. The paper is self-contained against external benchmarks, using pre-existing WMT annotations, published metrics, and standard meta-evaluation strategies; the protocol-mismatch and small-sample concerns are legitimate correctness risks, not circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new theoretical entities. Its central method relies on established human annotation protocols and meta-evaluation measures, plus one fitted tie-calibration threshold. The main burden is carried by domain assumptions about MQM as gold standard and the comparability of protocols.

free parameters (1)
  • tie-calibration threshold epsilon_e = estimated per evaluator
    Appendix C.2 defines acc*_eq with a threshold epsilon_e for each evaluator, estimated from the evaluator's score distribution. This fitted threshold determines which score differences count as ties and directly affects the acc*_eq rankings, where human baselines rank lower.
assumptions (4)
  • domain assumption MQM annotations released at WMT are a valid gold standard for translation quality.
    Section 2.2: 'Following standard practice in the literature ... we designate evaluators derived from the MQM annotations released annually at WMT as the ground truth.' All rankings are computed relative to this choice.
  • domain assumption Agreement with the gold-standard evaluator is the correct measure of evaluation quality for both metrics and humans.
    Section 2.3 defines SPA and acc*_eq as agreement with ground truth. The paper's own discussion in Section 4.1 questions whether higher agreement means better evaluation.
  • domain assumption Different human annotation protocols can be fairly compared on the same agreement yardstick.
    The paper compares ESA, pSQM, and DA+SQM baselines against MQM gold. Section 4 notes that acc*_eq may favor continuous metric scores over discrete human scores, so this fairness assumption is questionable.
  • domain assumption Restricting test sets to the largest subset of segments annotated by disjoint rater groups does not introduce selection bias.
    Section 2.1 solves an ILP to find this subset; footnote 3 acknowledges the 2023 EN->DE set drops to 145 segments, so the assumption may fail for that set.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress." pith.science (2026). https://pith.science/paper/HSBNYZIZ

@misc{pith2026250619571,
  author       = {Pith},
  title        = {Pith review of: Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HSBNYZIZ}},
  note         = {Machine review of arXiv:2506.19571}
}
read the original abstract

In Machine Translation (MT) evaluation, metric performance is assessed based on agreement with human judgments. In recent years, automatic metrics have demonstrated increasingly high levels of agreement with humans. To gain a clearer understanding of metric performance and establish an upper bound, we incorporate human baselines in the MT meta-evaluation, that is, the assessment of MT metrics' capabilities. Our results show that human annotators are not consistently superior to automatic metrics, with state-of-the-art metrics often ranking on par with or higher than human baselines. Despite these findings suggesting human parity, we discuss several reasons for caution. Finally, we explore the broader implications of our results for the research field, asking: Can we still reliably measure improvements in MT evaluation? With this work, we aim to shed light on the limits of our ability to measure progress in the field, fostering discussion on an issue that we believe is crucial to the entire MT evaluation community.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 37 canonical work pages

  1. [1]

    David Anugraha, Garry Kuwanto, Lucky Susanto, Derry Tanti Wijaya, and Genta Winata. 2024. https://doi.org/10.18653/v1/2024.wmt-1.32 M eta M etrics- MT : Tuning meta-metrics for machine translation via human preference calibration . In Proceedings of the Ninth Conference on Machine Translation, pages 459--469, Miami, Florida, USA. Association for Computati...

  2. [2]

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni. 2024. http...

  3. [3]

    Rachel Bawden, Biao Zhang, Andre T \"a ttar, and Matt Post. 2020. https://aclanthology.org/2020.wmt-1.98/ P ar BLEU : Augmenting metrics with automatic paraphrases for the WMT `20 metrics shared task . In Proceedings of the Fifth Conference on Machine Translation, pages 887--894, Online. Association for Computational Linguistics

  4. [4]

    Daniel Deutsch, Rotem Dror, and Dan Roth. 2021. https://doi.org/10.1162/tacl_a_00417 A statistical analysis of summarization evaluation metrics using resampling methods . Transactions of the Association for Computational Linguistics, 9:1132--1146

  5. [5]

    Daniel Deutsch, George Foster, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.798 Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12914--12929, Singapore. Association for Computational Linguistics

  6. [6]

    S \"o ren Dreano, Derek Molloy, and Noel Murphy. 2023 a . https://doi.org/10.18653/v1/2023.wmt-1.60 E mbed \_ L lama: Using LLM embeddings for the metrics shared task . In Proceedings of the Eighth Conference on Machine Translation, pages 738--745, Singapore. Association for Computational Linguistics

  7. [7]

    S \"o ren Dreano, Derek Molloy, and Noel Murphy. 2023 b . https://doi.org/10.18653/v1/2023.wmt-1.59 T okengram \_ F , a fast and accurate token-based chr F ++ derivative . In Proceedings of the Eighth Conference on Machine Translation, pages 730--737, Singapore. Association for Computational Linguistics

  8. [8]

    Muhammad ElNokrashy and Tom Kocmi. 2023. https://doi.org/10.18653/v1/2023.wmt-1.61 e BLEU : Unexpectedly good machine translation evaluation using simple word embeddings . In Proceedings of the Eighth Conference on Machine Translation, pages 746--750, Singapore. Association for Computational Linguistics

Show all 81 references
  1. [9]

    Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, Andr \'e Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat. 2023. https://doi.org/10.18653/v1/2023.wmt-1.100 The devil is in the errors: Leveraging large language models for f...

  2. [10]

    Mara Finkelstein, Dan Deutsch, Parker Riley, Juraj Juraska, Geza Kovacs, and Markus Freitag. 2024. https://arxiv.org/abs/2411.15387 From jack of all trades to master of one: Specializing llm-based autoraters to a test set . Preprint, arXiv:2411.15387

  3. [11]

    Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 a . https://doi.org/10.1162/tacl_a_00437 Experts, errors, and context: A large-scale study of human evaluation for machine translation . Transactions of the Association for C...

  4. [12]

    Markus Freitag, Nitika Mathur, Daniel Deutsch, Chi-Kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Frederic Blain, Tom Kocmi, Jiayi Wang, David Ifeoluwa Adelani, Marianna Buchicchio, Chrysoula Zerva, and Alon Lavie. 2024. https://doi.org/10.18653/v1/2024.wmt-1.2 Ar...

  5. [13]

    Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...

  6. [14]

    Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and Andr \'e F. T. Martins. 2022. https://aclanthology.org/2022.wmt-1.2/ Results of WMT 22 metrics shared task: Stop using BLEU -- neural metrics...

  7. [15]

    Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ond r ej Bojar. 2021 b . https://aclanthology.org/2021.wmt-1.73/ Results of the WMT 21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and n...

  8. [16]

    Thamme Gowda, Tom Kocmi, and Marcin Junczys-Dowmunt. 2023. https://doi.org/10.18653/v1/2023.wmt-1.62 Cometoid: Distilling strong reference-based machine translation metrics into E ven stronger quality estimation metrics . In Proceedings of the Eighth Conference on Machine Tran...

  9. [17]

    Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F

    Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F. T. Martins. 2024. https://doi.org/10.1162/tacl_a_00683 xcomet: Transparent machine translation evaluation through fine-grained error detection . Transactions of the Association for Co...

  10. [18]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . In International Conference on Learning Representations

  11. [19]

    Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. https://doi.org/10.18653/v1/2024.wmt-1.35 M etric X -24: The G oogle submission to the WMT 2024 metrics shared task . In Proceedings of the Ninth Conference on Machine Translation, pages 492--504, Miami...

  12. [20]

    Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle submission to the WMT 2023 metrics shared task . In Proceedings of the Eighth Conference on Machin...

  13. [21]

    Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.649 DEMETR : Diagnosing evaluation metrics for translation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Lang...

  14. [22]

    Fabio Kepler, Jonay Tr \'e nous, Marcos Treviso, Miguel Vera, and Andr \'e F. T. Martins. 2019. https://doi.org/10.18653/v1/P19-3020 O pen K iwi: An open source framework for quality estimation . In Proceedings of the 57th Annual Meeting of the Association for Computational Li...

  15. [23]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...

  16. [24]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, To...

  17. [25]

    Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...

  18. [26]

    Tom Kocmi and Christian Federmann. 2023 a . https://doi.org/10.18653/v1/2023.wmt-1.64 GEMBA - MQM : Detecting translation quality error spans with GPT -4 . In Proceedings of the Eighth Conference on Machine Translation, pages 768--775, Singapore. Association for Computational ...

  19. [27]

    Tom Kocmi and Christian Federmann. 2023 b . https://aclanthology.org/2023.eamt-1.19/ Large language models are state-of-the-art evaluators of translation quality . In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 193--203,...

  20. [28]

    Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. https://aclanthology.org/2021.wmt-1.57/ To ship or not to ship: An extensive evaluation of automatic metrics for machine translation . In Proceedings of the...

  21. [29]

    Tom Kocmi, Hitokazu Matsushita, and Christian Federmann. 2022 b . https://aclanthology.org/2022.wmt-1.47/ MS - COMET : More and better human judgements improve metric performance . In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 541--548, Abu Dhabi...

  22. [30]

    Tom Kocmi, Vil \'e m Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popovi \'c , Mrinmaya Sachan, and Mariya Shmatova. 2024 b . https://doi.org/10.18653/v1/2024.wmt-1.131 Error span annotation: A balanced approach for human evaluation of machine tra...

  23. [31]

    Samuel L \"a ubli, Rico Sennrich, and Martin Volk. 2018. https://doi.org/10.18653/v1/D18-1512 Has machine translation achieved human parity? a case for document-level evaluation . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages ...

  24. [32]

    Yilun Liu, Xiaosong Qiao, Zhanglin Wu, Su Chang, Min Zhang, Yanqing Zhao, Song Peng, Shimin Tao, Hao Yang, Ying Qin, Jiaxin Guo, Minghan Wang, Yinglu Li, Peng Li, and Xiaofeng Zhao. 2022. https://aclanthology.org/2022.wmt-1.48/ Partial could be better than whole. HW - TSC 2022...

  25. [33]

    Chi-kiu Lo. 2019. https://doi.org/10.18653/v1/W19-5358 Y i S i - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources . In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Pa...

  26. [34]

    Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 a . https://doi.org/10.5565/rev/tradumatica.77 Multidimensional quality metrics (mqm) : A framework for declaring and describing translation quality metrics . Tradumàtica, pages 455--463

  27. [35]

    Arle Richard Lommel, Maja Popovic, and Aljoscha Burchardt. 2014 b . Assessing inter-annotator agreement for translation error annotation. In MTE: Workshop on Automatic and Manual Metrics for Operational Translation Evaluation. International Conference on Language Resources and...

  28. [36]

    Federico Martelli, Stefano Perrella, Niccolò Campolungo, Tina Munda, Svetla Koeva, Carole Tiberius, and Roberto Navigli. 2025. https://doi.org/10.1162/coli_a_00541 Dibimt: A gold evaluation benchmark for studying lexical ambiguity in machine translation . Computational Linguis...

  29. [37]

    Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019. https://doi.org/10.18653/v1/P19-1269 Putting evaluation in context: Contextual embeddings improve machine translation evaluation . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics,...

  30. [38]

    Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ond r ej Bojar. 2020. https://aclanthology.org/2020.wmt-1.77/ Results of the WMT 20 metrics shared task . In Proceedings of the Fifth Conference on Machine Translation, pages 688--725, Online. Association for Computat...

  31. [39]

    Ananya Mukherjee, Hema Ala, Manish Shrivastava, and Dipti Misra Sharma. 2020. https://doi.org/10.1109/DSAA49011.2020.00042 Mee : An automatic metric for evaluation using embeddings for machine translation . In 2020 IEEE 7th International Conference on Data Science and Advanced...

  32. [40]

    Ananya Mukherjee and Manish Shrivastava. 2022 a . https://aclanthology.org/2022.wmt-1.50/ REUSE : RE ference-free U n S upervised quality estimation metric . In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 564--568, Abu Dhabi, United Arab Emirates ...

  33. [41]

    Ananya Mukherjee and Manish Shrivastava. 2022 b . https://aclanthology.org/2022.wmt-1.49/ Unsupervised embedding-based metric for MT evaluation with improved human correlation . In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 558--563, Abu Dhabi, U...

  34. [42]

    Ananya Mukherjee and Manish Shrivastava. 2023. https://doi.org/10.18653/v1/2023.wmt-1.66 MEE 4 and XL sim : IIIT HYD `s submissions' for WMT 23 metrics shared task . In Proceedings of the Eighth Conference on Machine Translation, pages 800--805, Singapore. Association for Comp...

  35. [43]

    Ananya Mukherjee and Manish Shrivastava. 2024. https://doi.org/10.18653/v1/2024.wmt-1.33 chr F - S : Semantics is all you need . In Proceedings of the Ninth Conference on Machine Translation, pages 470--474, Miami, Florida, USA. Association for Computational Linguistics

  36. [44]

    Subhajit Naskar, Daniel Deutsch, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.67 Quality estimation using minimum B ayes risk . In Proceedings of the Eighth Conference on Machine Translation, pages 806--811, Singapore. Association for Computational Linguistics

  37. [45]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...

  38. [46]

    Stefano Perrella, Lorenzo Proietti, Pere-Llu \'i s Huguet Cabot, Edoardo Barba, and Roberto Navigli. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.1152 Beyond correlation: Interpretable evaluation of machine translation metrics . In Proceedings of the 2024 Conference on...

  39. [47]

    Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Edoardo Barba, and Roberto Navigli. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.856 Guardians of the machine translation meta-evaluation: Sentinel metrics fall in! In Proceedings of the 62nd Annual Meeting of the...

  40. [48]

    Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Niccol \`o Campolungo, and Roberto Navigli. 2022. https://aclanthology.org/2022.wmt-1.51/ M a TES e: Machine translation evaluation as a sequence tagging problem . In Proceedings of the Seventh Conference on Machine Tra...

  41. [49]

    Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics

  42. [50]

    Maja Popovi \'c . 2017. https://doi.org/10.18653/v1/W17-4770 chr F ++: words helping character n-grams . In Proceedings of the Second Conference on Machine Translation, pages 612--618, Copenhagen, Denmark. Association for Computational Linguistics

  43. [51]

    Ricardo Rei, Jos \'e G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and Andr \'e F. T. Martins. 2022 a . https://aclanthology.org/2022.wmt-1.52/ COMET -22: Unbabel- IST 2022 submission for the metrics shared task . In...

  44. [52]

    Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, Andr \'e F. T. Martins, and Alon Lavie. 2021. https://aclanthology.org/2021.wmt-1.111/ Are references really needed? unbabel- IST 2021 submission for the metrics shared ...

  45. [53]

    Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G

    Ricardo Rei, Nuno M. Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 submission for the quality estimation shared t...

  46. [54]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 a . https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--270...

  47. [55]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 b . https://aclanthology.org/2020.wmt-1.101/ Unbabel`s participation in the WMT 20 metrics shared task . In Proceedings of the Fifth Conference on Machine Translation, pages 911--920, Online. Association for Compu...

  48. [56]

    Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G

    Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and Andr \'e F. T. Martins. 2022 b . https://aclanthology.org/2022.wmt-1.60/ C omet K iwi: IST -...

  49. [57]

    Parker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, and Markus Freitag. 2024. https://doi.org/10.18653/v1/2024.naacl-long.275 Finding replicable human evaluations via stable ranking probability . In Proceedings of the 2024 Conference of the North ...

  50. [58]

    Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 a . https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning robust metrics for text generation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. ...

  51. [59]

    Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur Parikh. 2020 b . https://aclanthology.org/2020.wmt-1.102/ Learning to evaluate translation beyond E nglish: BLEURT submissions to the WMT metrics 2020 shared task ....

  52. [60]

    Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. https://aclanthology.org/2006.amta-papers.25/ A study of translation edit rate with targeted human annotation . In Proceedings of the 7th Conference of the Association for Machine Translation...

  53. [61]

    Peter Stanchev, Weiyue Wang, and Hermann Ney. 2019. https://doi.org/10.18653/v1/W19-5359 EED : Extended edit distance measure for machine translation . In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 514--520, Florenc...

  54. [62]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...

  55. [63]

    Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Haji c , Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, and Roberto Navigli. 2023. https://doi.org/10.18653/v1/2023.acl-long.697 What`s the meaning of superhu...

  56. [64]

    Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024. https://doi.org/10.18653/v1/2024.wmt-1.118 Improving statistical significance in human evaluation of automatic metrics via soft pairwise accuracy . In Proceedings of the Ninth Conference on Machine Trans...

  57. [65]

    Brian Thompson and Matt Post. 2020 a . https://doi.org/10.18653/v1/2020.emnlp-main.8 Automatic machine translation evaluation in many languages via zero-shot paraphrasing . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages...

  58. [66]

    Brian Thompson and Matt Post. 2020 b . https://aclanthology.org/2020.wmt-1.67/ Paraphrase generation as zero-shot multilingual translation: Disentangling semantic similarity from lexical and syntactic diversity . In Proceedings of the Fifth Conference on Machine Translation, p...

  59. [67]

    Giorgos Vernikos, Brian Thompson, Prashant Mathur, and Marcello Federico. 2022. https://aclanthology.org/2022.wmt-1.6/ Embarrassingly easy document-level MT metrics: How to convert any pretrained metric into a document-level metric . In Proceedings of the Seventh Conference on...

  60. [68]

    Vasiliy Viskov, George Kokush, Daniil Larionov, Steffen Eger, and Alexander Panchenko. 2023. https://doi.org/10.18653/v1/2023.wmt-1.69 Semantically-informed regressive encoder score . In Proceedings of the Eighth Conference on Machine Translation, pages 815--821, Singapore. As...

  61. [69]

    Wong, Lidia S

    Yu Wan, Keqin Bao, Dayiheng Liu, Baosong Yang, Derek F. Wong, Lidia S. Chao, Wenqiang Lei, and Jun Xie. 2022 a . https://aclanthology.org/2022.wmt-1.53/ A libaba-translate C hina`s submission for WMT 2022 metrics shared task . In Proceedings of the Seventh Conference on Machin...

  62. [70]

    Yu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, and Lidia Chao. 2022 b . https://doi.org/10.18653/v1/2022.acl-long.558 U ni TE : Unified translation evaluation . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistic...

  63. [71]

    Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney. 2016. https://doi.org/10.18653/v1/W16-2342 C harac T er: Translation edit rate on character level . In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 505--510,...

  64. [72]

    Zhanglin Wu, Yilun Liu, Min Zhang, Xiaofeng Zhao, Junhao Zhu, Ming Zhu, Xiaosong Qiao, Jingfei Zhang, Ma Miaomiao, Zhao Yanqing, Song Peng, Shimin Tao, Hao Yang, and Yanfei Jiang. 2023. https://doi.org/10.18653/v1/2023.wmt-1.70 Empowering a metric with LLM -assisted named enti...

  65. [73]

    Jin Xu, Yinuo Guo, and Junfeng Hu. 2020. https://aclanthology.org/2020.wmt-1.104/ Incorporate semantic structures into machine translation evaluation via UCCA . In Proceedings of the Fifth Conference on Machine Translation, pages 934--939, Online. Association for Computational...

  66. [74]

    Wenda Xu, Xian Qian, Mingxuan Wang, Lei Li, and William Yang Wang. 2023. https://doi.org/10.18653/v1/2023.acl-long.283 SESCORE 2: Learning text generation evaluation via synthesizing realistic mistakes . In Proceedings of the 61st Annual Meeting of the Association for Computat...

  67. [75]

    Wenda Xu, Yi-Lin Tuan, Yujie Lu, Michael Saxon, Lei Li, and William Yang Wang. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.489 Not all errors are equal: Learning text generation metrics using stratified error synthesis . In Findings of the Association for Computation...

  68. [76]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...

  69. [77]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with bert . In Proceedings of the International Conference on Learning Representations

  70. [78]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://arxiv.org/abs/2306.05685 Judging llm-as-a-judge with mt-bench and chatbot arena . P...

  71. [79]

    Vil \'e m Zouhar, Shuoyang Ding, Anna Currey, Tatyana Badeka, Jenyuan Wang, and Brian Thompson. 2024. https://doi.org/10.18653/v1/2024.acl-short.45 Fine-tuned machine translation metrics struggle in unseen domains . In Proceedings of the 62nd Annual Meeting of the Association ...

  72. [80]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  73. [81]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.