REVIEW 12 cited by
xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xCOMET, an open-source learned metric designed to bridge the gap between these approaches. xCOMET integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xCOMET is largely capable of identifying localized critical errors and hallucinations.
Forward citations
Cited by 12 Pith papers
-
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
Using per-example in-context demonstrations built from historical same-source human MQM ratings makes an LLM judge dramatically better at fine-grained MT evaluation on WMT'23 and WMT'24.
-
Evaluating LLMs on Chinese Idiom Translation
Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).
-
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
A confidence-reward score for selecting preference pairs improves DPO-based machine translation fine-tuning over reward-only selection methods on ALMA-7B and NLLB-1.3B.
-
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
PromptOptMe compresses the inputs of the GEMBA-MQM MT evaluation prompt with a two-stage fine-tuned LLaMA 3.2 model, achieving a 2.37x token reduction without quality loss in the headline GPT-4o configuration.
-
Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation
Segment-level QE scores are not reliable enough to gate translation review, according to a new 104k-segment evaluation and a synthesis of prior critiques.
-
Hunyuan-MT Technical Report
Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.
-
Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
A 7B open-weight translation model matches or outperforms far larger commercial systems across 28 languages in automatic and human evaluations.
-
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
A new data-mixing recipe (Parallel-First Monolingual-Second) and a 9B model, GemmaX2-28, achieve translation quality competitive with Google Translate and GPT-4 across 28 languages.
-
MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation
MT-LENS is an open-source extension of LM-eval-harness that bundles MT quality, gender bias, added toxicity, and perturbation-robustness evaluations into one command-line and web-based toolkit.
-
A Measure of the System Dependence of Automated Metrics
A new metric, SysDep, quantifies how much an MT metric depends on the system being scored, and XCOMET's system dependence is large enough to change system rankings.
-
An Interdisciplinary Approach to Human-Centered Machine Translation
A position survey calling for human-centered machine translation, synthesizing translation studies and HCI to broaden MT evaluation and design beyond benchmark quality.
-
Reference-free Evaluation Metrics for Text Generation: A Survey
The survey classifies reference-free NLG evaluation metrics into learning from human judgments, pseudo-judgments, context-hypothesis correspondence, and peer evaluation.
Discussion (0). Continue with ORCID to comment.