{"id":"c0a73ba1-e072-47dd-b9eb-af1aa374629e","arxiv_id":"2508.18549","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Adding alternative translations as context improves automated MT quality scoring, with Kendall tau-b rising from 0.079 to up to 0.118.","lead":"This paper proposes two machine translation quality metrics that look beyond the single translation, using either other translations of the same sentence or retrieved translations with human scores. On segment-level tests, adding one extra candidate raises Kendall tau-b correlation with human judgment from 0.079 to about 0.118.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported gains could be label leakage through retrieved human scores in COMET-polyic; the abstract does not show that the retrieval pool is disjoint from the evaluation target.","rationale":"The reader's verdict is UNVERDICTED, and the specific weak point identified there is the most load-bearing condition: the provenance of COMET-polyic's retrieved labels. The abstract promises that 'incorporating retrieved examples in COMET-polyic yields similar improvements' but never states whether those examples come from the same benchmark used to compute tau-b. If they do, the metric can retrieve the human score of the very candidate being evaluated, and the gain is an artifact of the evaluation design rather than context-aware assessment. The same concern does not apply to COMET-polycand, which uses no labels, but the paper treats the two results as parallel support for one mechanism, so the second number is not optional. I also note that the supplied text is not machine-readable, so I could not confirm whether the full paper already documents a disjoint split; if it does, the concern would be resolved. In any case, the visible tables report point estimates only, and a paired bootstrap on the delta is necessary to support the word 'improves.' Public model release is a genuine reproducibility asset, and there is no evidence of bad faith; the issue is purely that the currently visible evidence cannot distinguish context-sensitive scoring from label leakage. Conditional acceptance is therefore the appropriate outcome: the central mechanism is plausible, but the leakage-free retrieval condition and an uncertainty interval for the delta must be demonstrated before the 0.079-to-0.116 gain is attributed to context-aware assessment.","tokens_in":18330,"tokens_out":9880,"duration_ms":102676,"concrete_test":"Re-run COMET-polyic on the same evaluation segments, but construct the retrieval index from a split made disjoint from the evaluation set (exclude all segment IDs/source sentences that appear in the target benchmark and use only human labels collected on segments outside that set), then recompute the segment-level Kendall tau-b. If the 0.116 value drops to the 0.079 baseline range, the headline gain is leakage-driven; if it is preserved, also report a paired bootstrap 95% CI for the delta so the 0.039-point improvement is distinguishable from noise.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that extra candidate context drives the improvement. For COMET-polyic the improvement (0.079 to 0.116 Kendall tau-b) is attributed to retrieved examples labeled with human quality scores. The abstract does not specify where those scores come from or whether the retrieval set is disjoint from the evaluation set. If the retrieval database includes the same source-candidate pairs that are being scored, the model receives the answer key as context; then the gain measures lookup, not context-aware assessment. The same concern does not apply to COMET-polycand, but the paper presents the two results as parallel evidence for one mechanism, so the second number is load-bearing too. The tables visible in the supplied text report only point estimates, so even a leakage-free reproduction would need a paired bootstrap CI on the 0.039-point delta to establish that it is not noise. Because the text is not fully machine-readable, I cannot rule out that §4 already specifies the split; if it does, this concern is settled.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two extensions of the COMET metric that condition on information beyond the standard single source-translation input. COMET-polycand scores a translation in the presence of alternative candidate translations of the same source sentence; COMET-polyic retrieves translations of similar source texts together with their human-labeled quality scores, following the retrieval-based in-context learning paradigm. The abstract reports that adding a single alternative candidate raises segment-level Kendall's tau-b correlation from 0.079 to 0.118, with further gains for more candidates, and that the retrieved-example variant reaches 0.116. The authors release the models publicly.","tokens_in":18484,"tokens_out":2373,"duration_ms":25244,"significance":"If the reported results are correct, the paper challenges the standard assumption that MT metrics should assess a translation in isolation; conditioning on extra candidate context would be a simple and potentially general improvement to existing learned metrics. The public release of the models is a valuable contribution and enables independent verification. The strength of the claim depends entirely on two conditions: that the gains are statistically reliable, and that the retrieved human-label information in COMET-polyic is not leaked from the evaluation target.","major_comments":[{"comment":"The central gain reported for COMET-polyic (0.079 to 0.116) is attributed to retrieved examples that carry human-labeled quality scores. The abstract does not state where these scores come from or whether the retrieval pool is disjoint from the segments being evaluated. If the retrieval index contains the same source-candidate pairs (or near-duplicates) that are scored at test time, then the model receives the answer key as part of its input and the gain measures lookup rather than context-aware assessment. The manuscript must specify the construction of the retrieval pool, its provenance, and the exact overlap-removal procedure used to keep evaluation segments out of the retrieval set; if Section 4 already does so, please point to the specific passage.","section":"Abstract; COMET-polyic description"},{"comment":"All reported improvements are point estimates without confidence intervals, significance tests, or run-level variance. For segment-level Kendall's tau-b, the difference between the baseline and the proposed methods (roughly 0.037-0.039) could easily lie within the noise band of a single evaluation set, particularly with a few thousand segments. The paper should provide paired bootstrap confidence intervals or an equivalent significance test over the evaluated test sets for the headline comparisons (baseline vs. polycand k=1, and baseline vs. polyic). Without such intervals, the claim that 'a single additional translation improves performance' is not statistically established.","section":"Abstract; Table 1 (k=1..8 rows)"},{"comment":"The abstract claims 'further gains when more translations are added,' which implies a monotone or at least increasing trend in the candidate count k. The visible table fragments do not allow verification of this trend, and it is possible that gains saturate or even reverse at larger k. The manuscript should report the full k-dependence explicitly and discuss whether the improvement is monotone; if it is not, the abstract's wording should be softened accordingly.","section":"Abstract; experimental tables"}],"minor_comments":[{"comment":"The baseline of 0.079 is not identified; please state which COMET variant and checkpoint are used as the single-translation baseline so that the comparison is reproducible.","section":"Abstract"},{"comment":"Please clarify whether the 'alternative translations' used as candidates are outputs from the same MT system as the translation being scored or from a pool of heterogeneous systems, since this affects the interpretation of the conditioning signal.","section":"COMET-polycand setup"},{"comment":"Please specify the provenance of the human quality scores for the retrieved examples (e.g., WMT human ratings, MQM labels), the number of score levels, and how the score distribution in the retrieval pool is matched to the evaluation distributions.","section":"COMET-polyic setup"}],"recommendation":"major_revision","confidential_remarks":"The supplied text of the manuscript is heavily corrupted by encoding issues, which prevented me from verifying whether Section 4 already addresses the retrieval-split concern. I recommend that the editor obtain a clean version before final decision, but the abstract-level ambiguity about label leakage and the absence of any statistical uncertainty in the reported headline numbers are sufficient to require revision. If the full paper already includes a clear disjointness guarantee and bootstrap intervals, the revision might be downgraded to minor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The idea is genuinely useful: scoring a translation after seeing other plausible translations of the same source is a natural input change, and the paper reports a real-looking jump in segment-level Kendall tau-b (0.079 to 0.118) from adding a single candidate, with more gains as candidates increase. And the authors release the models, which makes the claim checkable. That is the right kind of contribution.\n\nThe cleaner result is COMET-polycand. The alternative translations carry no human labels, so the gain cannot be label leakage. If it replicates, it is a practical, low-cost improvement to a widely used metric. The ICL variant, COMET-polyic, is where I would push hard. It feeds retrieved translations together with human quality scores into the input; the abstract does not say whether the retrieval pool is disjoint from the items being scored. If it is not, the 0.116 tau-b is partly lookup, not evaluation. That concern does not sink the paper, because the polycand result stands on its own, but the two numbers are presented as parallel evidence for one mechanism, so the authors need to state the split or run an ablation that removes the retrieved score labels.\n\nThe main weakness is reporting. The abstract gives point estimates only. There are no confidence intervals or significance tests on the 0.039 delta, and segment-level tau-b on WMT-style test sets can be noisy. A paired bootstrap or a per-segment variance estimate would settle whether the gain is real. This is a minor-to-moderate issue, not a reason to reject.\n\nI could not check the full text because the file I received is garbled, so novelty and related-work overlap are provisional. The abstract's framing does not obviously duplicate existing multi-reference or reranking metrics, but I cannot confirm the related-work section.\n\nWho this is for: MT evaluation people, especially the WMT metric track. A serious editor should send it to peer review. The empirical claim is concrete, the models are public, and the design is simple enough to replicate. I would want the revision to include error bars and a clear description of the retrieval split. I would be happy to referee it or have a student run the replication.","headline":"A modest, useful input-augmentation idea for COMET with a plausible headline gain, but the ICL variant needs a clean retrieval split and the abstract needs error bars before I'd trust the numbers.","tokens_in":19062,"tokens_out":3285,"would_cite":true,"duration_ms":33020,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-translation metrics match human judgment better when they rate a translation alongside alternative translations of the same source sentence.","keywords":["machine translation evaluation","COMET","Kendall tau-b","segment-level quality","in-context learning","retrieval-based scoring","candidate translations","human judgment correlation"],"falsifier":"Re-run COMET-polyic on a held-out set while replacing the human scores in the retrieved demonstrations with random or shuffled scores; if Kendall tau-b stays near 0.116, the model is using the score values rather than the source-translation content, and the claimed context effect would need reinterpretation.","tokens_in":18119,"feed_emoji":"🎯","tokens_out":5693,"duration_ms":59176,"temperature":0.7,"pith_summary":"Machine-translation metrics usually score a translation in isolation, whereas human assessment often involves comparing alternatives. The paper tries to close that gap by conditioning a learned metric on additional context. It proposes COMET-polycand, which adds alternative translations of the same source sentence, and COMET-polyic, which adds retrieved translations of similar sentences together with their human quality scores. Both variants raise segment-level agreement with human judgment: Kendall tau-b goes from 0.079 to 0.118 with one extra candidate, and to 0.116 with retrieved examples. If the result holds, evaluation of translation systems becomes more reliable without new annotation effort.","feed_headline":"One extra translation lifts MT metric correlation ~50%","feed_subtitle":"Kendall tau-b agreement with human quality judgments climbs from 0.079 to 0.118 with a single added candidate.","key_machinery":"The central machinery is the input constitution of the COMET cross-encoder, a model that reads a source sentence and a translation and produces a quality score. COMET-polycand extends that read: the source is accompanied not only by the candidate translation but by $k$ alternative translations of the same source, so the score is computed in a comparative setting. COMET-polyic extends it on the retrieval side: given a source segment, it finds similar source-translation pairs with known human scores and inserts them as labeled demonstrations, so the score is computed in an in-context-learning setting. The reported behavior, more context yielding more agreement with humans, is carried by this enlarged input rather than by a new network architecture.","core_discovery":"The paper's central discovery is that a learned machine-translation metric can be made more accurate by letting it score a translation in the presence of other translations rather than in isolation. COMET-polycand does this by feeding the source sentence, the candidate translation, and one or more alternative renderings of that same source into the encoder, turning quality assessment into a comparative judgment. COMET-polyic instead retrieves translations of similar source sentences plus their human quality scores and places them in the input as demonstrations, borrowing the retrieval-based in-context learning setup. The reported effect is a jump in segment-level Kendall tau-b correlation with human scores from 0.079 to 0.118 with a single additional candidate, further gains as candidates are added, and 0.116 for the retrieval variant. This is presented as evidence that the single-translation input assumption of standard metrics is a bottleneck.","pith_inferences":["Editorial extension: If the gains generalize, the standard practice of scoring each translation in isolation should be revisited across other learned metrics, since comparative context is likely a general feature of how humans assess quality.","Editorial extension: A direct stress test would shuffle the human labels in the retrieved demonstrations; if the improvement survives shuffling, the model is relying on label values rather than on the similarity structure.","Editorial extension: The comparative framing suggests a ranking-oriented training signal, where the metric is trained to order the candidate translations rather than to predict absolute scores, which is closer to how metrics are ultimately evaluated."],"forward_implications":["A single additional candidate translation already captures a large part of the total gain, so even cheaply available system outputs can improve metric accuracy.","Adding more candidates produces further improvements, so the metric can exploit existing multi-system evaluation runs without new human annotation.","Retrieval-based in-context examples provide a way to adapt the metric to a new domain or language pair on the fly, as long as labeled examples exist.","Because the models are released, other evaluators can adopt this input format without retraining from scratch."],"supporting_citations":[],"fun_headline_variants":["One extra translation lifts MT metric accuracy by 50%","MT metric correlation jumps 50% with a single extra candidate","Adding one translation boosts MT scoring context","MT evaluation improves with a second translation","COMET-poly: one extra candidate = 50% better correlation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The improvement assumes the extra translations and their human scores are not drawn from the same pool as the translation being judged, so the model cannot be reading the answer key through the additional input.","fun_headline_variants_meta":{"raw":{"variants":["One extra translation lifts MT metric accuracy by 50%","MT metric correlation jumps 50% with a single extra candidate","Adding one translation boosts MT scoring context","MT evaluation improves with a second translation","COMET-poly: one extra candidate = 50% better correlation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2340,"prompt_tokens":909,"completion_tokens":1431,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1354}},"tokens_in":525,"tokens_out":1431,"duration_ms":11166,"temperature":1.0,"reasoning_tokens":1354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:56:27.302968+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run COMET-polyic on a held-out set while replacing the human scores in the retrieved demonstrations with random or shuffled scores; if Kendall tau-b stays near 0.116, the model is using the score values rather than the source-translation content, and the claimed context effect would need reinterpretation.","supporting_citations":[],"review_version":2}