{"id":"8965e574-20b3-4182-9d4a-afa66ee95149","arxiv_id":"2508.14909","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Automatic metrics rank WMT25 MT systems, with human evaluation pending and expected to supersede these results.","lead":"This report gives preliminary machine translation quality rankings for the WMT25 shared task based on automatic metrics, with the caveat that human evaluation will later replace them. It is a service for participants writing their system papers, not a final verdict on translation quality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM judges GPT-4.1 and Command A are also evaluated systems; the paper never discloses this overlap, and self-preference in either judge would directly inflate its own AUTO RANK.","rationale":"The paper is genuinely careful: it labels results as preliminary, releases code and data, and lists multiple limitations. The central claim is modest — these are the rankings produced by a specified automatic procedure, to be superseded by human evaluation. But the claim that the rankings are 'preliminary findings' useful for system-description papers rests on the procedure being a fair reading of the metrics. A 40% weight on two LLM judges that are themselves contestants is the one unacknowledged structural bias. Unlike metric optimization, which the paper discloses, this overlap is not mentioned in Limitations. It is also directly testable now from released per-segment scores, so a conditional acceptance should require the test or a disclosure. This does not change the reader's CONDITIONAL verdict; it sharpens the condition.","tokens_in":43094,"tokens_out":7292,"duration_ms":78825,"concrete_test":"From the public repository, extract per-segment GEMBA-ESA scores for all LPs. For each LP: (1) z-normalize each judge's scores across systems; (2) compute Δ_GPT = mean(z_GPT4.1-judge on GPT-4.1 outputs) - mean(z_CMDA-judge on GPT-4.1 outputs), and Δ_CMDA analogously; (3) bootstrap 95% CIs over segments. If either Δ is significantly nonzero, self-preference is present. Additionally, re-run AUTO RANK excluding both GEMBA-ESA metrics; if GPT-4.1 or Command A moves more than 2 rank positions in any LP, the overlap changes the published ranking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The AUTO RANK gives 2 of 5 metrics to GEMBA-ESA judgments produced by GPT-4.1 and Command A, both of which are themselves in the evaluated system list (Table 1 and every per-LP table). For a rank ordering to be meaningful, these LLM judges must not systematically favor their own output style. The paper's Limitations section discusses metric optimization and unsupported languages but never mentions this judge-participant overlap. The GEMBA-ESA component is not blind: the judge model is evaluating translations that may share stylistic fingerprints with its own generation behavior. The released per-segment scores allow a direct check: for each LP, compare the score GPT-4.1-as-judge assigns to GPT-4.1-as-system against the score Command-A-as-judge assigns to the same system, after per-metric normalization; repeat for Command A. If either judge ranks itself above the other judge's ranking of it by more than the segment-level noise, the overlap is a real confound. The final AUTO RANK remaps z-scores linearly to 1..N, so even small mean shifts in GEMBA-ESA can move a system several rank positions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents preliminary rankings of WMT25 General Machine Translation systems, computed exclusively from automatic metrics and released to help participants write system description papers. The AUTO RANK procedure combines GEMBA-ESA with two LLM judges (GPT-4.1 and Command A), MetricX-24-Hybrid-XL, XCOMET-XL, and CometKiwi-XL (excluded for language pairs without human references), with chrF++ used alone for English-Bhojpuri and English-Maasai. System-level scores are averaged over ~100-word segments, normalized by median-interpercentile scaling (Eq. 1), averaged with equal weights (Eq. 2), and linearly rescaled to a 1..N rank. The authors explicitly warn that automatic rankings may favor re-ranking systems and that human evaluation will supersede them. Data, code, and per-segment scores are publicly released.","tokens_in":43373,"tokens_out":5693,"duration_ms":64304,"significance":"If the ranking is taken at face value, it provides a timely, transparent snapshot of system performance across 30+ language pairs and is a useful service to the WMT community. Strengths include the public release of test data, system outputs, per-segment scores, and LaTeX source; the aggregation procedure is clearly specified and easily reproducible; and the manuscript is unusually candid about metric bias, paragraph-level reliability, and low-resource limitations. However, the paper's central claim—that AUTO RANK produces a meaningful ordering—is weakened by an undisclosed judge-participant overlap (GPT-4.1 and Command A act both as GEMBA-ESA judges and as evaluated systems), by the absence of any uncertainty quantification, and by a tension between 'preliminary' framing and the statement that AUTO RANK is the official final result for systems not selected for human evaluation. These issues are testable with the released data and should be addressed before the ranking is used as a basis for system-description papers.","major_comments":[{"comment":"","section":"Automatic Ranking / Table 1"},{"comment":"No uncertainty quantification is provided for the reported ranks. Adjacent systems often differ by small fractions of an AutoRank unit (e.g., English-Czech rows: 3.5, 3.9, 4.4, 5.0; English-Russian rows: 4.2, 4.3, 4.4, 4.5). Given test sets of roughly 37K words per language pair and known segment-level metric noise, such differences may be statistically indistinguishable. The paper should provide bootstrap confidence intervals or a minimum-rank-difference threshold, and should avoid presenting one-decimal AUTO RANK values as if they represent a reliably ordered ranking. As written, the precision is misleading and the central claim of a meaningful ordering is not yet supported.","section":"Eq. (2) and per-LP tables"},{"comment":"The paper justifies equal-weight averaging by combining 'complementary failure modes,' but Appendix A shows that GEMBA-ESA-CMdA and GEMBA-ESA-GPT4.1 correlate above 0.8 for most language pairs (e.g., en-fa: 0.852; en-lt: 0.828; en-sv: 0.780), and the LLM-judge family is 2 of 5 metrics (2 of 4 when CometKiwi is excluded). Equal weighting therefore gives the LLM-judge family more effective weight than the nominal 2/5, and if the overlap concern above is real, the bias compounds. A robustness check—leave-one-metric-out or inverse-correlation reweighting—would show whether the ranking is stable under alternative aggregation schemes.","section":"Appendix A / Eq. (2)"},{"comment":"The Abstract and Introduction describe the rankings as 'preliminary,' but the Human Evaluation section states that 'for any system not selected for human evaluation, the automatic metric ranking (AUTO RANK) serves as the official final result.' Many systems in the per-LP tables are not marked for human evaluation (blank Humeval column), so a substantial portion of the reported ranking is final, not preliminary. This tension should be resolved: either the abstract should say that the automatic ranking is final for a subset of systems, or the official status should be qualified for those systems. This matters because finality raises the stakes of the judge-overlap and uncertainty issues discussed above.","section":"Human Evaluation"}],"minor_comments":[{"comment":"The sentence 'The official WMT25 ranking will be based on human evaluation, which is more reliable and will supersede these results.' is duplicated verbatim in both the abstract and the introduction.","section":"Abstract and Introduction"},{"comment":"Several rows contain the LaTeX artifact '— /times' (e.g., Claude-4, Gemini-2.5-Pro, GPT-4.1, Mistral-Medium); these should render as '—' or '✗' consistently.","section":"Table 1"},{"comment":"The repository link is malformed: '/githubgithub.com/wmt-conference/wmt-collect-translations' should be a proper URL.","section":"Prompts"},{"comment":"Minor typo: 'i.e.,Bhojpuri' should have a space after the comma. Also, the table header 'LP Supported' is ambiguous; consider renaming to 'LP officially supported by system' or similar.","section":"Low-resource exception"},{"comment":"Capitalization inconsistency: 'Gemba-ESA' appears in the Limitations section; the established acronym elsewhere is 'GEMBA-ESA'.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"This is a useful community resource with admirable data release and transparency. The dominant issue is the undisclosed judge-participant overlap, which is directly testable with the released per-segment scores; I would not accept the paper without that analysis or an explicit, argued justification for why self-preference is negligible. The uncertainty-quantification point is also important given that the automatic ranking is final for a large subset of systems. Both are fixable within the manuscript's scope, hence major revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a straightforward preliminary evaluation report for WMT25. The rankings are new for this year, the test sets and system outputs are public, and the aggregation scheme in Eqs. (1)-(2) is clearly specified and correctly implemented as far as I can tell from the tables. The paper is also unusually candid: it says outright that automatic rankings can be biased toward re-ranking, that human evaluation will supersede them, and that some participants optimized for GEMBA-ESA. That honesty is worth crediting.\n\nThe main thing I'd flag is the judge-participant overlap. GPT-4.1 and Command A serve as the two GEMBA-ESA judges, and both also appear as submitted systems in the same tables. The Limitations section discusses metric optimization and unsupported languages, but never mentions this. If either judge has any systematic self-preference, its own AUTO RANK position is inflated through two of the five metric slots. The released per-segment scores allow a direct check: compare how GPT-4.1-as-judge scores GPT-4.1-as-system versus how Command-A-as-judge scores it, and vice versa, after normalizing. That test should be run before anyone takes the fine-grained rankings seriously. The paper's own caveat that human evaluation will supersede these results softens the problem, but not fully: for any system not selected for human evaluation, the AUTO RANK is the official final result.\n\nThe other soft spot is the absence of uncertainty quantification. The paper gives a deterministic rank ordering with no confidence intervals or significance testing, and the final linear remapping to 1..N amplifies small mean shifts in GEMBA-ESA. That's a minor issue for a preliminary report, but it interacts with the judge-overlap problem.\n\nThe low-resource chrF++ decision is reasonable given the cited evidence, though the paper admits chrF++ correlates poorly with humans. The duplicated sentence in the abstract is a trivial editorial slip.\n\nOverall, the central claim – that these are preliminary, automatically computed rankings – holds up. The methodology is inherited, but the dataset and results are new. I'd send this to peer review with a request for the authors to disclose the overlap and either run the self-preference check or qualify the affected ranks. The paper is a useful community resource and the thinking is transparent.","headline":"A useful, honestly hedged preliminary WMT25 ranking that deserves review, but the undisclosed overlap between the LLM judges and the evaluated systems is a real confound, especially for systems whose automatic ranking becomes the official result.","tokens_in":43982,"tokens_out":2306,"would_cite":true,"duration_ms":25341,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Shy-hunyuan-MT tops WMT25's automatic ranking in every pair","keywords":["machine translation","WMT25","automatic evaluation","LLM-as-a-judge","quality estimation","AUTO RANK","GEMBA-ESA","shared task"],"falsifier":"Once WMT25 releases the human Error Span Annotation scores, compare the human ordering per language pair with AUTO RANK: if the two disagree badly for high-resource pairs (e.g., Spearman rank correlation below roughly 0.5), the claim that this metric ensemble is a meaningful quality ordering fails. A separate probe for the judge-participant overlap: have GPT-4.1 and Command A score a blinded set containing their own outputs under different system names and compare with scores from an independent judge; systematic self-inflation would invalidate their columns.","tokens_in":43019,"feed_emoji":"🏆","tokens_out":7403,"duration_ms":79239,"temperature":0.7,"pith_summary":"WMT25's General Machine Translation shared task has too many submissions to wait for human evaluation before participants start writing system descriptions, so this paper issues a preliminary ranking computed entirely by automatic metrics. The ranking method, named AUTO RANK, combines two reference-based metrics, one quality-estimation model, and two LLM judges, then rescales and averages them to produce one order per language pair. For the lowest-resource pairs, Bhojpuri and Maasai, the method falls back to the surface metric chrF++. The authors stress that the result is provisional: it may inflate systems that re-rank outputs with metrics or MBR, and the official WMT25 ranking, based on human Error Span Annotation, will replace it. Read as an early signal, the tables indicate which systems are worth studying; they are not final findings.","feed_headline":"Shy-hunyuan-MT tops WMT25's automatic ranking in every pair","feed_subtitle":"Metric-only table puts a 7B open model ahead of GPT-4.1 and Gemini until human judges deliver the final word.","key_machinery":"The central object is AUTO RANK, a named aggregation procedure. It converts each metric's system scores $x_s^{(m)}$ into robust z-scores $z_s^{(m)}=(x_s^{(m)}-\\tilde{x}^{(m)})/D^{(m)}$, where $\\tilde{x}^{(m)}$ is the median and $D^{(m)}=\\max(\\varepsilon,Q_{100}^{(m)}-Q_{25}^{(m)})$, averages the $z_s^{(m)}$ with equal weights to get $\\bar{z}_s$, and maps $\\{\\bar{z}_s\\}$ linearly onto $1,\\dots,N$ so lower is better. The construction is continuous and monotonic, so no system is discarded, within-metric order is preserved, extreme outliers do not dominate, and the relative spacing between systems survives the final remap. It is the piece of machinery that turns five heterogeneous metric columns","core_discovery":"The paper establishes the AUTO RANK protocol as the provisional official ranking mechanism for WMT25. For each typical language pair it scores every submitted system with GEMBA-ESA run by two LLM judges (GPT-4.1 and Command A), MetricX-24-Hybrid-XL, XCOMET-XL, and CometKiwi-XL, averaging paragraph-level scores and combining them by median-interpercentile scaling followed by an equal-weight mean and a linear remap to ranks 1 through N. Bhojpuri and Maasai, where learned metrics are unvalidated, are ranked by chrF++ alone. The resulting tables put Shy-hunyuan-MT, a 7B constrained system, first on every language pair shown, ahead of large proprietary models. The authors frame the output as prel","pith_inferences":["A blinded swap test could isolate judge self-bias: if GPT-4.1 and Command A are removed as judges and replaced by an independently judged LLM, their system ranks may shift, revealing whether judge-participant overlap inflates their positions.","Because the normalization keeps every system and preserves spacing, AUTO RANK is sensitive to which systems enter the pool; adding or removing a strong outlier changes everyone's scaled distance, so cross-language comparisons of rank gaps are not meaningful.","The low-correlation appendix suggests Marathi, whose learned metrics disagree most, is the weakest point of the multilingual subtrack; recomputing Marathi's ranking with chrF++ if references were available would test how much the order depends on metric choice."],"forward_implications":["If AUTO RANK reflects quality, participants can use these tables immediately to position their system descriptions while human evaluation is still running.","For the 16 language pairs without human evaluation, AUTO RANK is the final official ranking, so its reliability there is not merely preliminary.","Systems that use Quality Estimation or Minimum Bayes Risk re-ranking may be systematically over-ranked, so the tables should not be read as pure model-quality ordering.","The repeated first-place of a 7B constrained system implies small open models can beat much larger proprietary ones on this metric mix, a pattern the human results can confirm or refute."],"supporting_citations":[{"why":"Supplies GEMBA-ESA, the reference-less LLM-as-a-judge method used by both judges in AUTO RANK.","marker":"(Kocmi and Federmann, 2023)"},{"why":"Supplies MetricX-24-Hybrid-XL, one of the two trained reference-based metrics in the ensemble.","marker":"(Juraska et al., 2024)"},{"why":"Supplies XCOMET-XL, the other trained reference-based metric in the ensemble.","marker":"(Guerreiro et al., 2024)"},{"why":"Supplies CometKiwi-XL, the trained quality-estimation metric used for most language pairs.","marker":"(Rei et al., 2023)"},{"why":"Supplies chrF++, the surface-level metric used alone for Bhojpuri and Maasai.","marker":"(Popović, 2017)"},{"why":"Provides the evidence that reference-based and reference-less metrics have complementary failure modes, justifying the metric combination.","marker":"(Freitag et al., 2023)"},{"why":"Documents how systems can be optimized for metrics, grounding the paper's bias caveat about re-ranking.","marker":"(Finkelstein and Freitag, 2024)"},{"why":"Defines the Error Span Annotation protocol that will produce the human rankings superseding AUTO RANK.","marker":"(Kocmi et al., 2024)"}],"fun_headline_variants":["7B model leads WMT25 auto ranking, human judges pending","Metric-only WMT25 table puts 7B Shy-hunyuan first","WMT25 provisional: small model beats giants on auto metrics","Auto-rank WMT25: Shy-hunyuan-MT ahead until human eval","Prelim WMT25: 7B open model tops every language pair"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the five automatic metrics—including two LLM judges that are themselves among the ranked systems—behave like human quality judgments at paragraph level across all 30 language pairs, including low-resource pairs where they were never validated.","fun_headline_variants_meta":{"raw":{"variants":["7B model leads WMT25 auto ranking, human judges pending","Metric-only WMT25 table puts 7B Shy-hunyuan first","WMT25 provisional: small model beats giants on auto metrics","Auto-rank WMT25: Shy-hunyuan-MT ahead until human eval","Prelim WMT25: 7B open model tops every language pair"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1387,"prompt_tokens":673,"completion_tokens":714,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":624}},"tokens_in":417,"tokens_out":714,"duration_ms":7177,"temperature":1.0,"reasoning_tokens":624,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:34:11.967474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Once WMT25 releases the human Error Span Annotation scores, compare the human ordering per language pair with AUTO RANK: if the two disagree badly for high-resource pairs (e.g., Spearman rank correlation below roughly 0.5), the claim that this metric ensemble is a meaningful quality ordering fails. A separate probe for the judge-participant overlap: have GPT-4.1 and Command A score a blinded set containing their own outputs under different system names and compare with scores from an independent judge; systematic self-inflation would invalidate their columns.","supporting_citations":[],"review_version":1}