REVIEW 3 major objections 5 minor 50 references
MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read MT-LENS is an open-source toolkit that extends the LM-eval-harness to cover translation quality, gender bias, added toxicity, and robustness to character noise in one evaluation workflow.
desk verdict A useful MT evaluation toolkit that needs reference-implementation checks before I'd trust its numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the task abstraction inherited and extended from LM-eval-harness: each evaluation is a named task {src}_{tgt}_{dataset} whose dataset, prompt template, and metric configuration are declared in YAML, and whose results are emitted as JSON with segment-level scores. This lets MT-LENS treat quality, bias, toxicity, and noise robustness uniformly, and the UI layer then renders those JSON files as interactive comparisons. The bootstrapped t-test over BLEU, COMET, and COMET-KIWI is the significance machinery for system comparison.
What would settle it
Take a held-out set of English sentences with known correct and incorrect gendered translations (for example, nurse and doctor templates with the referent swapped), run the MUST-SHE and MMHB tasks through MT-LENS, and compare the reported accuracy against the known labels; if a deliberately wrong-gender translation is not flagged as incorrect, the gender-bias pipeline is not measuring what it claims.
Extended reading notes
Core claim
The paper introduces MT-LENS as a unified evaluation framework built on LM-eval-harness. It defines MT tasks by dataset plus language pair and supports five blocks: model backends (fairseq, CTranslate2, transformers, vllm, plus pre-generated translations), task definitions, prompt formatting, metrics, and JSON results. For quality, it supports BLEU, TER, CHRF, COMET, BLEURT, MetricX, XCOMET, COMET-KIWI, and quality-estimation variants; for added toxicity, it filters HOLISTIC BIAS source sentences with MUTOX and scores translations with ETOX, MUTOX, and DETOXIFY; for gender bias, it runs MUST-SHE, MMHB, and MT-GenEval tasks out of English; for robustness, it injects swap, character-duplication, and character-drop noise into FLORES-200 at a controllable level. The Streamlit interface shows error spans from XCOMET, segment-length scatter plots, bootstrapped significance tests, and per-task dashboards.
Load-bearing premise
The reliability of the gender-bias and added-toxicity scores rests entirely on the external classifiers and datasets that label toxicity and gender, and the paper does not check those labels against human judgments, so a mislabeling proxy would make the toolkit's outputs look valid while being wrong.
Editorial extensions
If this is right
- Users can run translation quality, gender bias, added toxicity, and character-noise robustness evaluations on the same generative model through a single command-line interface.
- The JSON output format with segment-level scores makes it possible to inspect exactly which sentences drive quality, bias, or toxicity differences between systems.
- Bootstrapped significance tests on BLEU, COMET, and COMET-KIWI let practitioners see whether observed system differences are statistically meaningful.
- Because it is built on LM-eval-harness, new MT datasets can be added by declaring tasks in YAML without changing evaluation code, and non-MT NLU tasks remain available in the same harness.
Reading between the lines
- If MT-LENS gains adoption, evaluation of MT systems could standardize around the same harness used for LLM benchmarks, making it easier to compare bias and toxicity results across papers—but only if the community agrees on which classifiers and thresholds to use.
- The paper's added-toxicity pipeline could be extended to non-English source languages by combining MUTOX's multilingual coverage with source-side toxicity classifiers, something the current HOLISTIC-BIAS-based setup only partially supports.
- The same perturbation framework could be used to test robustness to real-world OCR or keyboard noise if natural noise corpora were swapped in for the synthetic swap, chardupe, and chardrop operations.
- The UI's error-span visualization depends on XCOMET; a natural extension would be to let users click through to the underlying token-level scores or to compare spans produced by different error-detection models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MT-LENS is a proposed open-source toolkit that extends LM-eval-harness to machine translation evaluation, supporting translation quality metrics (BLEU, TER, CHRF, COMET, BLEURT, MetricX, XCOMET), three gender-bias tasks (MUST-SHE, MMHB, MT-GENEVAL), an added-toxicity pipeline (HOLISTIC BIAS with MUTOX/ETOX/DETOXIFY), and character-noise robustness perturbations. It also provides a Streamlit-based user interface for segment- and system-level analysis, including error-span visualization, segment-length scatter plots, and bootstrapped significance tests. The paper describes the architecture, lists supported datasets and metrics in tables, shows an example command, and reports qualitative UI observations from a Catalan-to-English comparison of madlad-400-3B and NLLB-3.3B.
Significance. If the toolkit works as described, it would be a useful contribution to the MT evaluation ecosystem: it builds on the widely adopted LM-eval-harness, unifies several evaluation tasks beyond translation quality, offers a user-friendly UI, and releases code on GitHub. The inclusion of bootstrapped significance tests and support for error-span visualization are concrete strengths. However, the paper's central claim—that MT-LENS provides reliable evaluation across these tasks—is not yet supported by any validation evidence; the described pipelines are asserted rather than checked against reference implementations or human judgments. Because the tool produces numeric scores that users will trust, the lack of correctness evidence is a load-bearing gap that must be addressed before the claims can be accepted.
major comments (3)
- [Section 3.2] The gender-bias and added-toxicity pipelines (MUST-SHE with the 'revised script of Mash et al. (2024)', MMHB via CHRF subset scores, MT-GENEVAL, and the MUTOX-filtered ETOX/MUTOX/DETOXIFY toxicity pipeline) are described but never validated against the original task scripts, an independent reference implementation, or human judgments. These are composite pipelines with many opportunities for off-by-one errors, morphological mismatches, incorrect subsetting, or classifier-threshold mistakes, so MT-LENS can report misleading bias and toxicity scores while running without errors. Please add a validation experiment that reproduces published dataset statistics or reference outputs on a small shared set (e.g., recompute MUST-SHE accuracy with the official script and compare, or verify toxicity labels against the original ETOX/MUTOX releases).
- [Section 3 and Table 2] The metric wrappers listed in Table 2 (SacreBLEU, unbabel-comet, metricx, transformers for BLEURT) are claimed to support 'state-of-the-art metrics', but the paper provides no evidence that MT-LENS reproduces the outputs of the underlying libraries. BLEU depends heavily on tokenization and smoothing choices, COMET versions differ in model weights and normalization, and MetricX has multiple variants, so a wrapper can silently produce different scores. Include a reproducibility check, such as running the same sentences through MT-LENS and the official implementations on a standard set (e.g., FLORES-200 devtest) and reporting score differences; this is essential for any evaluation tool.
- [Sections 1 and 4] The demo and demo-video links are placeholder text ('this link' in Section 1), and the UI demonstration in Section 4 reports only qualitative findings (e.g., 'madlad-400-3B exhibits greater robustness') without providing the underlying score tables, evaluation JSON, or a reproducible command sequence. This blocks independent verification of the central usability claim. Populate the links and include a small reproducible example with actual computed scores (e.g., BLEU/COMET values for a few segments) so readers can compare against their own runs.
minor comments (5)
- [Throughout] There are numerous spacing artifacts in key terms: the title uses 'MT-L ENS', Table 1 has 'H OLISTIC BIAS', 'M UST-SHE', and 'MT-G ENEVAL', and the example command in Section 3.1 contains '-- tr a n sl a tio n _ kw a r gs'; these should be corrected to 'MT-LENS', 'HOLISTIC BIAS', 'MUST-SHE', 'MT-GENEVAL', and '--translation_kwargs'.
- [Section 3.1] The example usage is not valid shell syntax; it mixes variable assignment with a JSON-like block and uses spaces within option names. Provide an actual command-line example that users can copy and run.
- [Section 2] The related work says 'MT-C OMPARE EVAL' and the CTranslate2 footnote reads 'CTranslate22'; these should be 'MT-ComparEval' and 'CTranslate2'.
- [Section 4.3] The gender-bias UI is described as having tabs for MUST-SHE and MMHB, but the paper does not explain how MT-GENEVAL results are visualized or integrated into the UI; clarify.
- [Tables 1 and 2] The paper does not specify which versions or splits of the datasets are used (e.g., HOLISTIC BIAS release, FLORES-200 devtest, NTREX-128 version); pinning these versions is important for reproducibility.
Circularity Check
No significant circularity: MT-LENS is an integration toolkit, not a derivation, so no claim reduces to its own inputs.
full rationale
MT-LENS is a systems/integration paper with no fitted parameters, no equations, and no predicted quantities; its central claim is that the toolkit unifies existing MT evaluation tasks and metrics. The metrics in Table 2 are wrappers around external libraries (SacreBLEU, unbabel-comet, metricx, transformers), and the tasks in Table 1 are standard datasets (FLORES-200, NTREX-128, Tatoeba, NTEU, Holistic Bias, MuST-SHE, MMHB, MT-GenEval). The only author-self-citations are transparent implementation choices: the MuST-SHE gender-accuracy measure is 'the revised script of (Mash et al., 2024)' and the added-toxicity pipeline follows García Gilabert et al. (2024) and Costa-jussà et al. (2024a). These are attributed inputs, not hidden outputs: the tool does not derive gender-bias scores from the same data used to fit them, and no 'prediction' is claimed from a fitted subset. The paper's actual weakness is unverified implementation correctness — no comparison against reference implementations or human judgments, and broken 'this link' placeholders for demo/video — but that is a validation/QA gap, not circularity. Therefore no circular step is exhibited.
Assumptions & free parameters
assumptions (5)
- domain assumption Automated metrics such as BLEU, COMET, and COMET-KIWI are valid proxies for translation quality.
- domain assumption The toxicity classifiers ETOX, MUTOX, and DETOXIFY correctly detect toxic content in translations across the evaluated languages.
- domain assumption The MUST-SHE revision script, MMHB placeholder-based CHRF groupings, and MT-GENEVAL counterfactuals provide valid measurements of gender bias.
- domain assumption XCOMET error spans accurately mark translation errors as critical, major, or minor.
- standard math The bootstrapped t-test with segment-level scores is a valid significance test for comparing MT systems.
Cite this review
Pith. "Pith review of MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation." pith.science (2026). https://pith.science/paper/QXTAAOJB
@misc{pith2026241211615,
author = {Pith},
title = {Pith review of: MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QXTAAOJB}},
note = {Machine review of arXiv:2412.11615}
}
read the original abstract
We introduce MT-LENS, a framework designed to evaluate Machine Translation (MT) systems across a variety of tasks, including translation quality, gender bias detection, added toxicity, and robustness to misspellings. While several toolkits have become very popular for benchmarking the capabilities of Large Language Models (LLMs), existing evaluation tools often lack the ability to thoroughly assess the diverse aspects of MT performance. MT-LENS addresses these limitations by extending the capabilities of LM-eval-harness for MT, supporting state-of-the-art datasets and a wide range of evaluation metrics. It also offers a user-friendly platform to compare systems and analyze translations with interactive visualizations. MT-LENS aims to broaden access to evaluation strategies that go beyond traditional translation quality evaluation, enabling researchers and engineers to better understand the performance of a NMT model and also easily measure system's biases.
Figures
Reference graph
Works this paper leans on
-
[1]
Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. 2024. https://arxiv.org/abs/2402.17733 Tower: An open multilingual large language model for translation-related tasks . Preprint, arXiv:2402.17733
arXiv 2024
-
[2]
Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, Min-Jae Hwang, Hirofumi Inaguma, Christopher Klaiber, Ilia Kulikov, Pengwei Li, Daniel Licht, Jean Maillard, Ruslan Mavlyutov, Alice Rakotoarison, Kaushik Ram Sadagopan, Abinesh Rama...
arXiv 2023
-
[3]
Yonatan Belinkov and Yonatan Bisk. 2018. Synthetic and natural noise both break neural machine translation. In Proceedings of the International Conference on Learning Representations
work page 2018
-
[4]
Di Gangi, Roldano Cattoni, and Marco Turchi
Luisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia A. Di Gangi, Roldano Cattoni, and Marco Turchi. 2020. https://doi.org/10.18653/v1/2020.acl-main.619 Gender in danger? evaluating speech translation technology on the M u ST - SHE corpus . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6923--6933, On...
-
[5]
Laurent Bi \'e , Aleix Cerd \`a -i Cuc \'o , Hans Degroote, Amando Estela, Mercedes Garc \' a-Mart \' nez, Manuel Herranz, Alejandro Kohan, Maite Melero, Tony O ' Dowd, Sin \'e ad O ' Gorman, M \=a rcis Pinnis, Roberts Rozis, Riccardo Superbo, and Art \=u rs Vasi l evskis. 2020. https://aclanthology.org/2020.eamt-1.60 Neural translation for the E uropean ...
work page 2020
-
[6]
Marta Costa-juss \`a , David Dale, Maha Elbayad, and Bokai Yu. 2024 a . https://aclanthology.org/2024.eamt-1.31 Added toxicity mitigation at inference time for multimodal and massively multilingual translation . In Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1), pages 360--372, Sheffield, UK. Europ...
work page 2024
-
[7]
Marta Costa-juss \`a , Mariano Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, and Carleigh Wood. 2024 b . https://doi.org/10.18653/v1/2024.findings-acl.340 M u T ox: Universal MU ltilingual audio-based TOX icity dataset and zero-shot detector . In Findings of the Association for Computatio...
-
[8]
Marta Costa-juss \`a , Eric Smith, Christophe Ropers, Daniel Licht, Jean Maillard, Javier Ferrando, and Carlos Escolano. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.642 Toxicity in multilingual machine translation at scale . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9570--9586, Singapore. Association for Com...
Show all 50 references
-
[9]
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansan...
2022 arXiv
-
[10]
Anna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu, Benjamin Hsu, and Georgiana Dinu. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.288 MT - G en E val: A counterfactual and contextual dataset for evaluating gender accuracy in mac...
2022 doi
-
[11]
Christian Federmann, Tom Kocmi, and Ying Xin. 2022. https://doi.org/10.18653/v1/2022.sumeval-1.4 NTREX -128 -- news test references for MT evaluation of 128 languages . In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 21--24, Online. Associatio...
2022 doi
-
[12]
Batya Friedman and Helen Nissenbaum. 1995. https://doi.org/10.1145/223355.223780 Minimizing bias in computer systems . In Human Factors in Computing Systems, CHI '95 Conference Companion: Mosaic of Creativity, Denver, Colorado, USA, May 7-11, 1995 , page 444. ACM
1995
-
[13]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[14]
Javier Garc \' a Gilabert, Carlos Escolano, and Marta Costa-juss \`a . 2024. https://aclanthology.org/2024.eamt-1.8 R e S e TOX : Re-learning attention weights for toxicity mitigation in machine translation . In Proceedings of the 25th Annual Conference of the European Associa...
2024
-
[15]
Guerreiro, Pierre Colombo, Pablo Piantanida, and Andr \'e Martins
Nuno M. Guerreiro, Pierre Colombo, Pablo Piantanida, and Andr \'e Martins. 2023 a . https://doi.org/10.18653/v1/2023.acl-long.770 Optimal transport for unsupervised hallucination detection in neural machine translation . In Proceedings of the 61st Annual Meeting of the Associa...
2023 doi
-
[16]
Nuno M Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e FT Martins. 2023 b . xcomet: Transparent machine translation evaluation through fine-grained error detection. arXiv preprint arXiv:2310.10482
2023 arXiv
-
[17]
Laura Hanu and Unitary team . 2020. Detoxify. Github. https://github.com/unitaryai/detoxify
2020
-
[18]
Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 MetricX-23: The Google Submission to the WMT 2023 Metrics Shared Task . In Proceedings of the Eighth Conference on Machine Tr...
2023 doi
-
[19]
Ondrej Klejch, Eleftherios Avramidis, Aljoscha Burchardt, and Martin Popel. 2015. MT-ComparEval : Graphical evaluation interface for Machine Translation development. The Prague Bulletin of Mathematical Linguistics, 104:63--74
2015
-
[20]
Philipp Koehn. 2004. https://aclanthology.org/W04-3250 Statistical significance tests for machine translation evaluation . In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 388--395, Barcelona, Spain. Association for Computational...
2004
-
[21]
Philipp Koehn and Rebecca Knowles. 2017. https://doi.org/10.18653/v1/W17-3204 Six challenges for neural machine translation . In Proceedings of the First Workshop on Neural Machine Translation, pages 28--39, Vancouver. Association for Computational Linguistics
2017 doi
-
[22]
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2024. Madlad-400: A multilingual and document-level large audited dataset. Advances in Neural Information Processing Systems, 36
2024
-
[23]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...
2023
-
[24]
Audrey Mash, Carlos Escolano, Aleix Sant, Maite Melero, and Francesca de Luca Fornaciari. 2024. https://aclanthology.org/2024.lrec-main.1489 Unmasking biases: Exploring gender bias in E nglish- C atalan machine translation through tokenization analysis and novel dataset . In P...
2024
-
[25]
Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, and Xinyi Wang. 2019. https://doi.org/10.18653/v1/N19-4007 compare-mt: A tool for holistic comparison of language generation systems . In Proceedings of the 2019 Conference of the North A merican Chapter of the A...
2019 doi
-
[26]
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations
2019
-
[27]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...
2002
-
[28]
Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Edoardo Barba, and Roberto Navigli. 2024. https://doi.org/10.18653/v1/2024.acl-long.856 Guardians of the machine translation meta-evaluation: Sentinel metrics fall in! In Proceedings of the 62nd Annual Meeting of the As...
2024 doi
-
[29]
Ben Peters and Andr \'e FT Martins. 2024. Did translation models get more robust without anyone even noticing? arXiv preprint arXiv:2403.03923
2024
-
[30]
Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics
2015 doi
-
[31]
Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics
2018 doi
-
[32]
Ricardo Rei, Jos \'e G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and Andr \'e F. T. Martins. 2022 a . https://aclanthology.org/2022.wmt-1.52 COMET -22: Unbabel- IST 2022 submission for the metrics shared task . In ...
2022
-
[33]
Ricardo Rei, Ana C Farinha, Craig Stewart, Luisa Coheur, and Alon Lavie. 2021. https://doi.org/10.18653/v1/2021.acl-demo.9 MT - T elescope: A n interactive platform for contrastive evaluation of MT systems . In Proceedings of the 59th Annual Meeting of the Association for Comp...
2021 doi
-
[34]
Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and Andr \'e F. T. Martins. 2022 b . https://aclanthology.org/2022.wmt-1.60 C omet K iwi: IST -u...
2022
-
[35]
Aleix Sant, Carlos Escolano, Audrey Mash, Francesca De Luca Fornaciari, and Maite Melero. 2024. https://doi.org/10.18653/v1/2024.gebnlp-1.7 The power of prompts: Evaluating and mitigating gender bias in MT with LLM s . In Proceedings of the 5th Workshop on Gender Bias in Natur...
2024 doi
-
[36]
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof-Arenas, and Luisa Bentivogli. 2024 a . https://aclanthology.org/2024.emnlp-main.1002 What the harm? quantifying the tangible impact of gender bias in machine translation with a human-centered study . In Proceedings of th...
2024
-
[37]
Beatrice Savoldi, Andrea Piergentili, Dennis Fucci, Matteo Negri, and Luisa Bentivogli. 2024 b . https://aclanthology.org/2024.eacl-short.23 A prompt response to the demand for automatic gender-neutral translation . In Proceedings of the 18th Conference of the European Chapter...
2024
-
[38]
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning robust metrics for text generation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. Ass...
2020 doi
-
[39]
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.625 `` I ' m sorry to hear that '' : Finding new biases in language models with a holistic descriptor dataset . In Proceedings of the 202...
2022 doi
-
[40]
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. https://aclanthology.org/2006.amta-papers.25 A study of translation edit rate with targeted human annotation . In Proceedings of the 7th Conference of the Association for Machine Translation ...
2006
-
[41]
Craig Stewart, Ricardo Rei, Catarina Farinha, and Alon Lavie. 2020. https://aclanthology.org/2020.amta-user.4 COMET - deploying a new state-of-the-art MT evaluation metric in production . In Proceedings of the 14th Conference of the Association for Machine Translation in the A...
2020
-
[42]
Costa-jussà
Xiaoqing Ellen Tan, Prangthip Hansanti, Carleigh Wood, Bokai Yu, Christophe Ropers, and Marta R. Costa-jussà. 2024. https://arxiv.org/abs/2407.00486 Towards massive multilingual holistic bias . Preprint, arXiv:2407.00486
2024 arXiv
-
[43]
J \"o rg Tiedemann. 2020. https://aclanthology.org/2020.wmt-1.139 The tatoeba translation challenge -- realistic data sets for low resource and multilingual MT . In Proceedings of the Fifth Conference on Machine Translation, pages 1174--1182, Online. Association for Computatio...
2020
-
[44]
Jonas-Dario Troles and Ute Schmid. 2021. https://aclanthology.org/2021.wmt-1.61 Extending challenge sets to uncover gender bias in machine translation: Impact of stereotypical verbs and adjectives . In Proceedings of the Sixth Conference on Machine Translation, pages 531--541,...
2021
-
[45]
Bram Vanroy, Arda Tezcan, and Lieve Macken. 2023. https://aclanthology.org/2023.eamt-1.52 MATEO : MA chine translation evaluation online . In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 499--500, Tampere, Finland. Europe...
2023
-
[46]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020 doi
-
[47]
Emmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, and Andr \'e FT Martins. 2024. Watching the watchers: Exposing gender disparities in machine translation quality estimation. arXiv preprint arXiv:2410.10995
2024 arXiv
-
[48]
Vil \'e m Zouhar, Pinzhen Chen, Tsz Kin Lam, Nikita Moghe, and Barry Haddow. 2024. https://doi.org/10.18653/v1/2024.wmt-1.121 Pitfalls and outlooks in using COMET . In Proceedings of the Ninth Conference on Machine Translation, pages 1272--1288, Miami, Florida, USA. Associatio...
2024 doi
-
[49]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[50]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.