Pith. sign in

REVIEW 6 major objections 6 minor 60 references

TransBench: Benchmarking Machine Translation for Industrial-Scale Applications

T0 review · 6 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read TransBench tests machine translation for e-commerce use, scoring 17,000 sentences on domain knowledge, cultural fit, and robustness.

desk verdict A sensible framework for industrial MT evaluation wrapped around benchmark claims that are unverifiable and arithmetically wrong; not ready for review. read the letter →

arxiv 2505.14244 v1 pith:RSBMMHV6 submitted 2025-05-20 cs.CL

classification cs.CL
keywords machinetranslatione-commercebenchmarkevaluationmetricsqualityestimationculturaladaptationhallucinationmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that general-purpose MT benchmarks mislead industrial users, because e-commerce translation is judged by domain terminology, cultural appropriateness, and resilience to noisy input, not by fluency alone. To address this, it proposes a three-level capability framework — Basic Linguistic Competence, Domain-Specific Proficiency, and Cultural Adaptation — and introduces TransBench, a dataset of 17,000 professionally translated e-commerce sentences across four scenarios and 33 language pairs. TransBench pairs standard metrics (BLEU, TER, chrF, COMET) with Marco-MOS, a domain-specific quality-estimation model reported to reach 0.65 Pearson correlation with human scores. If the benchmark holds up, teams building commercial translation systems gain a way to compare models on the failures that actually occur online.

What carries the argument

The load-bearing machinery is a three-level evaluation grid that maps each capability level to dedicated data and metrics: Basic Linguistic Competence is probed by general metrics such as BLEU, TER, and chrF plus hallucination-rate and robustness BLEU-drop tests; Domain-Specific Proficiency is scored on scenario datasets using Marco-MOS, a quality-estimation model fine-tuned on 35,000 e-commerce MOS annotations; and Cultural Adaptation is measured by accuracy on honorific and taboo-word test items. This grid is what turns a collection of sentences into a diagnostic instrument rather than a simple leaderboard.

What would settle it

Take a random sample of TransBench sentences, have two independent teams of professional translators produce references, and compare BLEU and Marco-MOS scores computed against each reference; if model rankings flip or the difference between reference sets is large, the benchmark's reference-quality premise fails. Separately, re-running the stated 1,900-sentence Marco-MOS test set against fresh human MOS scores should reproduce a Pearson correlation near 0.65 if the headline result is stable.

Watch

Extended reading notes

Core claim

The paper presents TransBench as the first publicly available benchmark built specifically for e-commerce machine translation, grounded in a three-level capability framework that separates basic linguistic accuracy, domain-specific proficiency, and cultural adaptation. The dataset contains 17,000 professionally translated and verified sentences from four real e-commerce scenarios spanning 33 language pairs, alongside evaluation sets for hallucination, robustness, honorifics, and taboo words. The central quantitative claim is that Marco-MOS, a domain-tuned quality-estimation model, predicts human translation-quality judgments better than generic metrics or generic LLM evaluators, with a reported Pearson correlation of 0.65 and relative gains of +47.7% over GPT-4, +57.9% over COMET XL, and +64.98% over COMET XXL. The paper's claim is that industrial MT evaluation should be organized by capability level rather than by a single generic quality score.

Load-bearing premise

The benchmark's value rests on the assumption that the professionally produced reference translations are correct and consistent, but the paper reports no agreement scores between translators and does not yet release the data for inspection.

Editorial extensions

If this is right

  • E-commerce MT models can be compared on scenario-specific data across many language pairs, exposing weaknesses in product-listing, marketing, customer-service, and review translation that generic news benchmarks miss.
  • A domain-tuned quality-estimation model such as Marco-MOS gives practitioners a cheaper proxy for human evaluation when tuning e-commerce translation systems.
  • Hallucination rate and robustness BLEU-drop scores provide deployable guardrails: models that fail these tests can be filtered before release.
  • The same construction recipe, applied to finance and law data, would yield analogous industrial benchmarks rather than one universal test.
  • Scores on TransBench will not necessarily agree with WMT rankings, since cultural fidelity and domain terminology are weighted as first-class dimensions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the three-level framework generalizes, the future of MT evaluation is a family of domain-specific suites, each with its own quality-estimation model, rather than a single universal leaderboard.
  • The reported Marco-MOS advantage suggests that a modest amount of in-domain MOS data, on the order of tens of thousands of sentences, can outperform much larger generic evaluators, a tradeoff worth testing in finance and legal domains.
  • The hallucination and robustness subsets could be reused as stress tests for general-purpose LLM translation, independent of e-commerce, because those failure modes are not domain-specific.
  • A natural checkpoint for any user of TransBench is to re-score a sample with independently produced references, since no inter-annotator agreement figure appears in the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 6 minor

Summary. TransBench proposes a three-level framework (Basic Linguistic Competence, Domain-Specific Proficiency, Culture Adaptation) for evaluating industrial machine translation, introduces a benchmark called TransBench supposedly containing 17,000 professionally translated sentences over 33 language pairs for e-commerce translation, and reports a domain-specific quality estimation model called Marco-MOS with a claimed Pearson correlation of 0.65 with human judgments. The paper promises a publicly available benchmark and open-sourced evaluation tools, but it provides no release location, no evaluation results for any translation system on the benchmark, and its reported numerical claims contain internal inconsistencies. The manuscript is therefore more of a framework proposal and metric description than a usable benchmark release.

Significance. If TransBench and Marco-MOS were actually released and externally validated, the resource could fill a real gap in e-commerce MT evaluation, and the three-level framework is a reasonable organizing principle for reporting strengths and weaknesses across linguistic, domain, and cultural dimensions. However, the paper's central contribution is not verifiable in its current form: the dataset is not accessible, the dataset statistics are contradictory (17k sentences and 33 language pairs versus 16 languages and 60 language pairs plus additional data categories in Table 1), and the reported Marco-MOS improvements over GPT-4, COMET XL, and COMET XXL do not follow arithmetically from the quoted correlations. No system-level benchmark results are presented, so the paper cannot be judged as an evaluation resource. These issues are load-bearing rather than presentation-level, and they undermine the paper's headline claims.

major comments (6)
  1. [Section 4.2.1 and Multilingual Dimension paragraph] The abstract and Section 4.2.1 describe TransBench as containing 17,000 sentences over 33 language pairs, but the 'Multiligual Dimension' paragraph lists 16 languages and 60 language pairs, while Table 1 adds separate datasets of 12k, 2.6k, 29k, 232, and 107 items for finance, robustness, hallucination, taboo, and honorific evaluation. The relationship between these numbers is never explained: it is unclear whether 60 pairs is the planned scope, whether the 17k count includes only the e-commerce subset, or how Table 1's categories fit into the total. The paper must reconcile these figures, since the size and coverage of the benchmark are central to its claimed contribution.
  2. [Section 5.3] The reported improvement percentages are arithmetically inconsistent with the quoted correlations. If Marco-MOS achieves a Pearson correlation of 0.65 and GPT-4 achieves 0.4765, the relative improvement is (0.65-0.4765)/0.4765 = 36.4%, not +47.7%; against COMET XL (0.4267) the improvement is 52.3%, not +57.9%; and against COMET XXL (0.4457) it is 45.8%, not +64.98%. At least one set of numbers is wrong, so the state-of-the-art claim for Marco-MOS is not supported as stated and must be recomputed and corrected.
  3. [Abstract, Contributions, and Section 6] The paper repeatedly claims that TransBench is 'the first publicly available benchmark for e-commerce translation' and that evaluation tools are 'open-sourced,' but no repository URL, download link, or supplementary archive is provided anywhere in the manuscript. A benchmark contribution cannot be verified without access to the artifact; the paper needs an explicit release location and, ideally, a data card with licensing and access instructions before this claim can be evaluated.
  4. [Sections 4-5] The manuscript contains no evaluation results for any machine translation system on TransBench. There are no tables or figures reporting BLEU, TER, chrF, hallucination rates, robustness BLEU drops, taboo accuracy, or honorific accuracy for any model, and no comparison of representative MT systems on the proposed datasets. The only quantitative result is the Marco-MOS correlation in Section 5.3. Without such experiments, the paper does not demonstrate that the benchmark is usable, that its metrics separate models along the claimed dimensions, or that practitioners can compare systems against baseline scores.
  5. [Section 5.3 and Section 5.1] Marco-MOS is trained and tested exclusively on the authors' own manually annotated MOS data, with no inter-annotator agreement measure and no external validation set, so the 'state-of-the-art in E-commerce domain' claim rests on a self-constructed comparison. Additionally, Section 5.1 states that Marco-MOS is used 'for the financial datasets,' while Section 5.3 reports its performance in the e-commerce domain; the intended domain of the metric should be clarified. The metric should be validated on an independent, publicly available data set with established human judgments.
  6. [Section 4.2.2] The annotation pipeline is described only at the level of process steps ('double-blind cross-verification' and a '1-5 scale' scoring system), with no inter-annotator agreement scores, no per-language-pair statistics, and no quality-control metrics. Since the reference translations produced by this pipeline are the basis for every BLEU, TER, and Marco-MOS score computed on TransBench, the absence of annotation reliability evidence makes the benchmark's reference quality unverifiable and should be addressed with concrete agreement and quality numbers.
minor comments (6)
  1. [Section 4.2.2] The heading 'Mannual Annotation pipeline' contains a typo ('Mannual' should be 'Manual').
  2. [Section 4.1] The text says 'Table?? demonstrates the details of the TransBench'; the table reference is unresolved and should be fixed.
  3. [Throughout] 'Evaluation Matrices' is used where 'Evaluation Metrics' is meant, in the title of Section 5, in Figure 1, and elsewhere; this should be corrected throughout.
  4. [Multilingual Dimension paragraph] The language pair 'hd-en' appears in the list; given the language list includes Hindi, this should presumably be 'hi-en'.
  5. [Equations (1)-(4)] The definitions are sloppy: 'H_data represents the hallucination rate translation dataset' is not a meaningful definition, F(S_H|S) is introduced but not precisely specified, and Equations (2)-(4) reuse hallucination terminology for taboo and honorific data. These formulas should be rewritten with clear set notation.
  6. [Figure 8 caption] The caption says 'the reference text is the translated text using honorifics' for what is described as taboo data; this appears to be a copy-paste error and should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: TransBench and Marco-MOS are constructed artifacts, not derivations that reduce to their own inputs.

full rationale

No load-bearing step in the paper equates an output with an input by definition or by self-citation. The three-level framework is explicitly grounded in external translation theories (Catford 1965, Skopos theory, Cultural Turn), and the benchmark data are described as products of an annotation pipeline rather than derived from the evaluation metrics. Marco-MOS is a supervised quality-estimation model: the paper reports that its training and test data are human MOS scores (35k train / 1.9k test), and measuring Pearson correlation against held-out human scores is the standard operating procedure for such metrics, not a prediction forced by construction. The related-work citations by co-authors (e.g., Wang et al. 2023, Wu et al. 2024) are background support for LLM translation capability and are not load-bearing for the benchmark or metric claims. The paper's genuine weaknesses are external-validity and internal-consistency issues, not circularity: the claimed +47.7%/+57.9%/+64.98% improvements do not recompute from the quoted baseline correlations, the dataset is described as both 33 and 60 language pairs, and no release URL is provided for the 'publicly available' benchmark. These concerns should be raised as verifiability and correctness risks, but they do not constitute a circular derivation chain.

Assumptions & free parameters 4 free parameters · 4 assumptions · 3 invented entities

The central claims depend entirely on the authors' own construction and self-reported numbers. The benchmark and metric are not released, the human annotation process is not quantitatively validated, and the evaluation of Marco-MOS uses only the authors' test data, leaving no external checks.

free parameters (4)
  • Marco-MOS fine-tuning hyperparameters = 5 epochs, learning rate 5e-5, warmup 200 steps
    Chosen without reported ablations or sensitivity analysis; the claimed 0.65 correlation depends on these settings.
  • Hallucination detection thresholds = unspecified
    Section 5.2.3 defines hallucination via 'excessively large or small' vector distance and length-ratio thresholds without giving values.
  • Taboo and honorific matching lists = not released
    Accuracy formulas in Section 5.4 depend on lists of taboo words and honorific units that are not included in the paper.
  • Dataset balancing ratio = 35k training / 1.9k test samples
    The authors state the dataset was 'balanced' to these sizes, but the balancing procedure and distribution are not described.
assumptions (4)
  • domain assumption Human MOS scores are a reliable ground truth for translation quality.
    Marco-MOS is trained and evaluated against expert MOS ratings; the paper does not discuss subjectivity or annotation noise.
  • ad hoc to paper The three-level framework (linguistic, domain, cultural) is a complete decomposition of industrial MT quality.
    Proposed in Section 3 without empirical justification that the levels are distinct or exhaustive.
  • domain assumption Manual translation with double-blind verification yields high-quality references.
    Section 4.2.2 describes the process but reports no inter-annotator agreement or quality metrics.
  • domain assumption The selected e-commerce scenarios are representative of international e-commerce text.
    Four scenarios are listed, but no sampling statistics or coverage analysis are provided.
invented entities (3)
  • TransBench benchmark dataset
    purpose: Evaluate MT models on e-commerce translation across three capability levels.
    The dataset is claimed to be publicly available, but no URL or release exists in the paper, so it cannot be independently checked.
  • Marco-MOS
    purpose: Automated quality scoring model for e-commerce MT.
    Performance is only reported on the authors' own test set; no code, model weights, or external evaluation are provided.
  • F-MOS
    purpose: Planned finance-domain quality scoring model listed in Table 1.
    Mentioned only as future work; not implemented or described further.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TransBench: Benchmarking Machine Translation for Industrial-Scale Applications." pith.science (2026). https://pith.science/paper/RSBMMHV6

@misc{pith2026250514244,
  author       = {Pith},
  title        = {Pith review of: TransBench: Benchmarking Machine Translation for Industrial-Scale Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RSBMMHV6}},
  note         = {Machine review of arXiv:2505.14244}
}
read the original abstract

Machine translation (MT) has become indispensable for cross-border communication in globalized industries like e-commerce, finance, and legal services, with recent advancements in large language models (LLMs) significantly enhancing translation quality. However, applying general-purpose MT models to industrial scenarios reveals critical limitations due to domain-specific terminology, cultural nuances, and stylistic conventions absent in generic benchmarks. Existing evaluation frameworks inadequately assess performance in specialized contexts, creating a gap between academic benchmarks and real-world efficacy. To address this, we propose a three-level translation capability framework: (1) Basic Linguistic Competence, (2) Domain-Specific Proficiency, and (3) Cultural Adaptation, emphasizing the need for holistic evaluation across these dimensions. We introduce TransBench, a benchmark tailored for industrial MT, initially targeting international e-commerce with 17,000 professionally translated sentences spanning 4 main scenarios and 33 language pairs. TransBench integrates traditional metrics (BLEU, TER) with Marco-MOS, a domain-specific evaluation model, and provides guidelines for reproducible benchmark construction. Our contributions include: (1) a structured framework for industrial MT evaluation, (2) the first publicly available benchmark for e-commerce translation, (3) novel metrics probing multi-level translation quality, and (4) open-sourced evaluation tools. This work bridges the evaluation gap, enabling researchers and practitioners to systematically assess and enhance MT systems for industry-specific needs.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 32 canonical work pages

  1. [1]

    Nature, 630 0 (8018): 0 841--846, 2024

    Scaling neural machine translation to 200 languages. Nature, 630 0 (8018): 0 841--846, 2024

  2. [2]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  3. [3]

    D. M. Alves, J. Pombal, N. M. Guerreiro, P. H. Martins, J. Alves, M. A. Farajian, B. Peters, R. Rei, P. Fernandes, S. Agrawal, P. Colombo, J. G. C. de Souza, and A. F. T. Martins. Tower: An open multilingual large language model for translation-related tasks. CoRR, abs/2402.17733, 2024. doi:10.48550/ARXIV.2402.17733. URL https://doi.org/10.48550/arXiv.2402.17733

  4. [4]

    TICO-19: the Translation Initiative for Covid-19

    A. Anastasopoulos, A. Cattelan, Z.-Y. Dou, M. Federico, C. Federman, D. Genzel, F. Guzm \'a n, J. Hu, M. Hughes, P. Koehn, et al. Tico-19: the translation initiative for covid-19. arXiv preprint arXiv:2007.01788, 2020

  5. [5]

    Bahdanau, K

    D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. In Y. Bengio and Y. LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , 2015. URL http://arxiv.org/abs/1409.0473

  6. [6]

    Banerjee and A

    S. Banerjee and A. Lavie. METEOR : An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , pages 65--72, Ann Arbor, Michigan, June 2005. Association for Computational Linguistics. URL https://www.ac...

  7. [7]

    Bassnett and A

    S. Bassnett and A. Lefevere, editors. Translation, history and culture. Pinter Publishers, 1990

  8. [8]

    Caswell, E

    I. Caswell, E. Nielsen, J. Luo, C. Cherry, G. Kovacs, H. Shemtov, P. Talukdar, D. Tewari, B. M. Diane, K. M. Doumbouya, et al. Smol: Professionally translated parallel data for 115 under-represented languages. arXiv preprint arXiv:2502.12301, 2025

Show all 60 references
  1. [9]

    J. C. Catford. A linguistic theory of translation: an essay in applied linguistics. Oxford University Press, 1965

  2. [10]

    K. Cho, B. van Merri \"e nboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. Learning phrase representations using RNN encoder -- decoder for statistical machine translation. In A. Moschitti, B. Pang, and W. Daelemans, editors, Proceedings of the 2014 Conf...

  3. [11]

    Colombo, N

    P. Colombo, N. Guerreiro, R. Rei, D. Van, L. Coheur, and A. Martins. xcomet: Transparent machine translation evaluation through fine-grained error detection. Transactions of the Association for Computational Linguistics, 2023

  4. [12]

    Communication, L

    S. Communication, L. Barrault, Y. Chung, M. C. Meglioli, D. Dale, N. Dong, P. Duquenne, H. Elsahar, H. Gong, K. Heffernan, J. Hoffman, C. Klaiber, P. Li, D. Licht, J. Maillard, A. Rakotoarison, K. R. Sadagopan, G. Wenzek, E. Ye, B. Akula, P. Chen, N. E. Hachem, B. Ellis, G. M....

  5. [13]

    M. R. Costa - juss \` a , J. Cross, O. C elebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Y. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe...

  6. [14]

    Deutsch, E

    D. Deutsch, E. Briakou, I. Caswell, M. Finkelstein, R. Galor, J. Juraska, G. Kovacs, A. Lui, R. Rei, J. Riesa, et al. Wmt24++: Expanding the language coverage of wmt24 to 55 languages & dialects. arXiv preprint arXiv:2502.12404, 2025

  7. [15]

    A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El - Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, M. Auli, and A. Joulin. Beyond english-centric multilingual machine translation. J. Mach. Learn. Res., 22: 0 107:1--10...

  8. [16]

    Gehring, M

    J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin. Convolutional sequence to sequence learning. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 , volume ...

  9. [17]

    Goyal, C

    N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzm \'a n, and A. Fan. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10: 0 522...

  10. [18]

    Graham, T

    Y. Graham, T. Baldwin, A. Moffat, and J. Zobel. Continuous measurement scales in human evaluation of machine translation. In A. Pareja-Lora, M. Liakata, and S. Dipper, editors, Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 33-...

  11. [19]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  12. [20]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  13. [21]

    Haddow, T

    B. Haddow, T. Kocmi, P. Koehn, and C. Monz, editors. Proceedings of the Ninth Conference on Machine Translation, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. doi:10.18653/v1/2024.wmt-1.0. URL https://aclanthology.org/2024.wmt-1.0/

  14. [22]

    Hatim and I

    B. Hatim and I. Mason. Discourse and the translator. Longman, 1990

  15. [23]

    Hearne and A

    M. Hearne and A. Way. Statistical machine translation: A guide for linguists and translators. Lang. Linguistics Compass, 5 0 (5): 0 205--226, 2011. doi:10.1111/J.1749-818X.2011.00274.X. URL https://doi.org/10.1111/j.1749-818X.2011.00274.x

  16. [24]

    Hurst, A

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024

  17. [25]

    Hutchins

    J. Hutchins. Latest developments in machine translation technology: Beginning a new era in MT research. In Proceedings of Machine Translation Summit IV, pages 11--34, Kobe, Japan, July 19-22 1993. URL https://aclanthology.org/1993.mtsummit-1.2/

  18. [26]

    Jiang, T

    Y. Jiang, T. Liu, S. Ma, D. Zhang, J. Yang, H. Huang, R. Sennrich, R. Cotterell, M. Sachan, and M. Zhou. BlonDe : An automatic evaluation metric for document-level machine translation. In M. Carpuat, M.-C. de Marneffe, and I. V. Meza Ruiz, editors, Proceedings of the 2022 Conf...

  19. [27]

    Johnson, M

    M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Vi \'e gas, M. Wattenberg, G. Corrado, M. Hughes, and J. Dean. G oogle`s multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computationa...

  20. [28]

    Juraska, D

    J. Juraska, D. Deutsch, M. Finkelstein, and M. Freitag. M etric X -24: The G oogle submission to the WMT 2024 metrics shared task. In B. Haddow, T. Kocmi, P. Koehn, and C. Monz, editors, Proceedings of the Ninth Conference on Machine Translation, pages 492--504, Miami, Florida...

  21. [29]

    K. Knight. Building a large ontology for machine translation. In H uman L anguage T echnology: Proceedings of a Workshop Held at Plainsboro, New Jersey, March 21-24, 1993 , 1993. URL https://aclanthology.org/H93-1036/

  22. [30]

    Knight and S

    K. Knight and S. K. Luk. Building a large-scale knowledge base for machine translation. In B. Hayes - Roth and R. E. Korf, editors, Proceedings of the 12th National Conference on Artificial Intelligence, Seattle, WA, USA, July 31 - August 4, 1994, Volume 1, pages 773--778. AAA...

  23. [31]

    Kocmi and C

    T. Kocmi and C. Federmann. GEMBA - MQM : Detecting translation quality error spans with GPT -4. In P. Koehn, B. Haddow, T. Kocmi, and C. Monz, editors, Proceedings of the Eighth Conference on Machine Translation, pages 768--775, Singapore, Dec. 2023 a . Association for Computa...

  24. [32]

    Kocmi and C

    T. Kocmi and C. Federmann. Large language models are state-of-the-art evaluators of translation quality. In M. Nurminen, J. Brenner, M. Koponen, S. Latomaa, M. Mikhailov, F. Schierl, T. Ranasinghe, E. Vanmassenhove, S. A. Vidal, N. Aranberri, M. Nunziatini, C. P. Escart \'i n,...

  25. [33]

    P. Koehn. Statistical Machine Translation. Cambridge University Press, 2010. ISBN 978-0-521-87415-1. URL http://www.statmt.org/book/

  26. [34]

    A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024

  27. [35]

    A. R. Lommel, A. Burchardt, and H. Uszkoreit. Multidimensional quality metrics: a flexible system for assessing translation quality. In Proceedings of Translating and the Computer 35, London, UK, Nov. 28-29 2013. Aslib. URL https://aclanthology.org/2013.tc-1.6/

  28. [36]

    D. W. Lonsdale, T. Mitamura, and E. Nyberg. Acquisition of large lexicons for practical knowledge-based MT . Mach. Transl., 9 0 (3-4): 0 251--283, 1994. doi:10.1007/BF00980580. URL https://doi.org/10.1007/BF00980580

  29. [37]

    Lopes, M

    A. Lopes, M. A. Farajian, R. Bawden, M. Zhang, and A. F. T. Martins. Document-level neural MT : A systematic comparison. In A. Martins, H. Moniz, S. Fumega, B. Martins, F. Batista, L. Coheur, C. Parra, I. Trancoso, M. Turchi, A. Bisazza, J. Moorkens, A. Guerberof, M. Nurminen,...

  30. [38]

    A. Lopez. Statistical machine translation. ACM Comput. Surv. , 40 0 (3): 0 8:1--8:49, 2008. doi:10.1145/1380584.1380586. URL https://doi.org/10.1145/1380584.1380586

  31. [39]

    Michel and G

    P. Michel and G. Neubig. Mtnt: A testbed for machine translation of noisy text. arXiv preprint arXiv:1809.00388, 2018

  32. [40]

    M \"u ller, A

    M. M \"u ller, A. Rios, E. Voita, and R. Sennrich. A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation. In O. Bojar, R. Chatterjee, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, C. Monz, ...

  33. [41]

    E. A. Nida. Toward a science of translating: with special reference to principles and procedures involved in Bible translating. Brill, 1964

  34. [42]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. B leu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--318, Philadelphia, Pennsylvania, USA, July 2002 a . Associati...

  35. [43]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311--318, 2002 b

  36. [44]

    Popovi \'c

    M. Popovi \'c . chrf: character n-gram f-score for automatic mt evaluation. In Proceedings of the tenth workshop on statistical machine translation, pages 392--395, 2015

  37. [45]

    Przybocki, K

    M. Przybocki, K. Peterson, S. Bronsart, and G. Sanders. The nist 2008 metrics for machine translation challenge—overview, methodology, metrics, and results. Machine Translation, 23: 0 71--103, 2009

  38. [46]

    R. Rei, C. Stewart, A. C. Farinha, and A. Lavie. Comet: A neural framework for mt evaluation. arXiv preprint arXiv:2009.09025, 2020

  39. [47]

    R. Rei, A. C. Farinha, C. Zerva, D. van Stigt, C. Stewart, P. Ramos, T. Glushkova, A. F. T. Martins, and A. Lavie. Are references really needed? unbabel- IST 2021 submission for the metrics shared task. In L. Barrault, O. Bojar, F. Bougares, R. Chatterjee, M. R. Costa-jussa, C...

  40. [48]

    Rei and H

    K. Rei and H. J. Vermeer. Grundlegung einer allgemeine Translationstheorie. Max Niemeyer Verlag, 1984

  41. [49]

    Salesky, M

    E. Salesky, M. Federico, and M. Carpuat, editors. Proceedings of the 21st International Conference on Spoken Language Translation (IWSLT 2024), Bangkok, Thailand (in-person and online), Aug. 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.iwslt-1.0/

  42. [50]

    Sellam, D

    T. Sellam, D. Das, and A. Parikh. BLEURT : Learning robust metrics for text generation. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online, July 2020...

  43. [51]

    Snover, B

    M. Snover, B. Dorr, R. Schwartz, L. Micciulla, and J. Makhoul. A study of translation edit rate with targeted human annotation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers, pages 223--231, Cambridge, Massach...

  44. [52]

    Sutskever, O

    I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Pro...

  45. [53]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Pro...

  46. [54]

    Vernikos, B

    G. Vernikos, B. Thompson, P. Mathur, and M. Federico. Embarrassingly easy document-level MT metrics: How to convert any pretrained metric into a document-level metric. In P. Koehn, L. Barrault, O. Bojar, F. Bougares, R. Chatterjee, M. R. Costa-juss \`a , C. Federmann, M. Fishe...

  47. [55]

    L. Wang, C. Lyu, T. Ji, Z. Zhang, D. Yu, S. Shi, and Z. Tu. Document-level machine translation with large language models. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 16646--16661, ...

  48. [56]

    M. Wu, Y. Yuan, G. Haffari, and L. Wang. (perhaps) beyond human translation: Harnessing multi-agent collaboration for translating ultra-long literary texts. CoRR, abs/2405.11804, 2024. doi:10.48550/ARXIV.2405.11804. URL https://doi.org/10.48550/arXiv.2405.11804

  49. [57]

    Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnic...

  50. [58]

    H. Xu, Y. J. Kim, A. Sharaf, and H. H. Awadalla. A paradigm shift in machine translation: Boosting translation performance of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net...

  51. [59]

    H. Xu, A. Sharaf, Y. Chen, W. Tan, L. Shen, B. V. Durme, K. Murray, and Y. J. Kim. Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, Ju...

  52. [60]

    Zhang, V

    T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.