{"id":"7dee0285-1599-4404-9c21-1ba5e5e12798","arxiv_id":"2412.20127","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"M-MAD decouples MQM criteria into four dimensions and uses per-dimension multi-agent debate, achieving better WMT23 meta-evaluation scores than prior LLM-as-a-judge methods and rivaling learned metrics.","lead":"This paper introduces M-MAD, a multi-agent LLM framework that splits machine translation evaluation into four dimensions, holds a debate between two agents per dimension, and merges the results into an MQM-style score. On the WMT23 benchmark, the method beats other LLM-based judges and approaches top learned metrics, even when powered by the small model GPT-4o mini.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hyperparameters (debate topic, strategy, rounds) are selected on the WMT23 ZH-EN test set and the headline meta-score includes that same set; no held-out validation or significance estimate is provided, so the reported gains may be a selection artifact.","rationale":"The reader's weakest assumption is the same load-bearing concern I identify: the configuration (Severity, Consensus, R=3) was selected on the WMT23 ZH-EN test set, and the headline result includes that set. If this assumption fails, the reported ZH-EN meta-score and the Table 2 average are inflated, and the claimed superiority over GEMBA-MQM and EAPrompt may not reproduce. The paper provides no held-out validation or variance estimate; because temperature is set to 0, the LLM is deterministic, so resampling the model is not an option, and the only meaningful uncertainty is exactly the test-set/config-selection interaction that is unquantified. The ablation in Table 3 strengthens the concern indirectly: removing the debate stage changes the meta-score by only -0.006, meaning the method's distinctive component is not the main driver of performance, which makes the margin over baselines look even less stable. I am not arguing the result is false: the EN-DE results and HE-EN results in Appendix D suggest transferability, and the matched-backend comparison against GEMBA-MQM and EAPrompt is fair. The paper deserves a conditional acceptance rather than outright rejection, but a single nested split-half configuration-selection experiment on ZH-EN would settle whether the headline is an artifact. If that test passes, the conditional verdict could be upgraded; if it fails, the central claim would need to be substantially weakened. For these reasons, I agree with the reader's conditional assessment and recommend no change to the verdict.","tokens_in":21689,"tokens_out":15778,"duration_ms":158127,"concrete_test":"Split the WMT23 ZH-EN test set into two halves by system (e.g., 8 systems for tuning, 9 for evaluation); rerun the configuration search from Tables 6, 7, and Figure 3 on the tuning half, freeze the winning topic/strategy/round count, and compute the M-MAD meta-score on the evaluation half. Repeat with swapped halves and, if feasible, 10 random system-level splits. If the average held-out meta-score for M-MAD does not exceed GEMBA-MQM and EAPrompt on the same held-out halves, or if the margin over GEMBA-MQM drops below the reported 0.030, then the ZH-EN headline is a test-set selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that M-MAD outperforms all compared LLM-as-a-judge methods rests on a configuration search performed on the evaluation set itself. Tables 6, 7, and Figure 3 select the debate topic (Severity), debate strategy (Consensus), and number of rounds (R=3) by maximizing meta-evaluation scores on WMT23 ZH-EN, and the same ZH-EN results are then included in the headline average and abstract claims. The reported ZH-EN score is therefore the maximum over the explored configurations, not the performance of a pre-specified method. No held-out validation split, standard-error estimate, or significance test is provided, and with temperature 0 the only uncertainty is exactly the test-set-and-configuration interaction that is left unquantified. The EN-DE results are encouraging and provide partial evidence of transfer, but the headline average still mixes the tuned ZH-EN result with the untuned EN-DE result. In addition, the ablation in Table 3 shows that removing the debate stage (Stage 2) changes the ZH-EN meta by only -0.006, so the named multi-agent-debate component contributes little; this does not refute the empirical claim, but it underscores the fragility of the reported margin.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes M-MAD, a three-stage LLM-based framework for machine translation evaluation. Stage 1 decomposes the MQM annotation guideline into four dimensions (accuracy, fluency, style, terminology); Stage 2 runs two-agent 'Pro-Con' debates within each dimension, with a consensus checker; Stage 3 synthesizes dimension-level viewpoints into a final MQM-style score. Using GPT-4o mini at temperature 0, the authors report WMT23 meta-evaluation results on ZH-EN, EN-DE, and HE-EN, claiming that M-MAD outperforms existing LLM-as-a-judge methods and is competitive with state-of-the-art learned automatic metrics, despite being training-free and reference-free. Detailed ablations identify dimension partition as the largest contributor, and appendix analyses discuss error-span prediction and cases of possibly mislabeled gold annotations.","tokens_in":21938,"tokens_out":8945,"duration_ms":86058,"significance":"If the reported results are robust, the paper would make a useful contribution: it shows that a prompt-based, training-free LLM judge can approach the segment-level performance of learned metrics such as MetricX-23 and XCOMET on WMT23, and it provides a systematic decomposition of the MQM task. The strengths include matched GPT-4o mini backends across LLM-as-a-judge baselines, use of the standard MTME evaluation tool, publicly available code and data, and a transparent ablation structure. However, the significance is tempered by the fact that key framework choices (debate topic, debate strategy, number of rounds) were selected by maximizing the same ZH-EN test-set meta-score that is then included in the headline results, and by the very small measured contribution of the multi-agent debate component itself.","major_comments":[{"comment":"The debate topic (Severity), debate strategy (Consensus), and maximum debate rounds (R=3) are selected by maximizing meta-evaluation scores on the WMT23 ZH-EN test set, and the same ZH-EN scores are then included in the Table 2 average and in the abstract claims. Table 6 shows large swings across topics (0.808 vs. 0.666 vs. 0.737), Table 7 shows smaller but still positive differences for Consensus, and Figure 3 selects R=3. This makes the reported ZH-EN result a maximum over searched configurations rather than the performance of a pre-specified method. The EN-DE results provide partial evidence of transfer, but the headline average still mixes the tuned ZH-EN result with the untuned EN-DE result. Please provide a held-out validation procedure (for example, tuning on one language pair and reporting only the other pair as the test set, or using WMT22 for configuration selection) and clearly separate tuned from untuned results.","section":"Section 4.3, Tables 6-7, Figure 3"},{"comment":"No significance estimates or confidence intervals are reported anywhere, and the margins that carry the framework-level claims are extremely small. The ablation in Table 3 shows that removing Stage 2 (the multi-agent debate) changes the ZH-EN meta score by only -0.006; Table 7 shows the Consensus strategy at 0.808 versus the no-debate baseline at 0.802. Because temperature is 0, these are deterministic outputs, but the underlying test set and configuration interaction are not quantified. A difference of 0.006 in a composite meta-score cannot be distinguished from noise without bootstrap or other uncertainty estimates. Please add such estimates, or explicitly identify which conclusions remain supported without them.","section":"Tables 3 and 7, Section 3.2"},{"comment":"The appendix states that 'M-MAD is also superior among all metrics in the He-En task,' but Table 9 shows M-MAD with a meta score of 0.787, below MetricX-23 (0.807), XCOMET-QE-Ensemble (0.789), and MaTESe (0.792), with segment-level Pearson also below several learned metrics. This factual contradiction undermines the claim of consistent cross-lingual superiority and should be corrected. It also matters for the abstract's 'competes with state-of-the-art reference-based automatic metrics' claim, since the HE-EN results are weaker than the ZH-EN/EN-DE results.","section":"Appendix D, Table 9"},{"comment":"The paper's title and central narrative emphasize the multi-agent debate component, but the decoupled multidimensional design (Stage 1) is what drives most of the improvement (-0.041 meta when removed), while removing Stage 2 changes the meta score by only -0.006. This is not a refutation of the empirical headline, but it is a mismatch between the framework's claimed mechanism and the evidence. Please report more granular results for the debate stage (for example, per-dimension accuracy, error-span F1 changes, or case-level agreement) so that the contribution of Stage 2 is characterized substantively rather than through a single near-zero delta.","section":"Table 3 and Section 2.2"}],"minor_comments":[{"comment":"The text says 'In the reference-based setting, M-MAD surpasses COMETKiwi by 2.6% and MetricX-23-QE by 0.9%,' but COMETKiwi and MetricX-23-QE are reference-free metrics. The sentence should refer to the reference-free setting.","section":"Section 3.2"},{"comment":"The caption says 'Metrics with gray background is reference-based.' The verb should agree, and it would help to clarify that the gray background is not visible in the text-only rendering.","section":"Table 1 caption"},{"comment":"The second Style few-shot example annotates the span 'merchants' in a translation that does not contain the word 'merchants.' This appears to be a copy-paste error and may degrade the few-shot prompt for the Style agent.","section":"Figure 13"},{"comment":"The claim that system-level performance 'consistently peaks at round 3' is not supported by error bars or quantitative values in the figure. Please state the exact values and, if possible, add variability estimates.","section":"Section 4.4, Figure 3"},{"comment":"The caption is ambiguous about the first numerical column; it should state explicitly that the META column is the average meta-evaluation score across ZH-EN and EN-DE, while the subsequent columns are per-language component scores.","section":"Table 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The skeptic's main concern is justified: the framework configuration is selected on the WMT23 ZH-EN test set and the same set appears in the headline average. The small effect of the debate stage and the contradiction in the HE-EN appendix reinforce the need for a revised presentation with held-out validation and uncertainty quantification. I would accept a revision that addresses these points; the underlying idea is publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the empirical core is honest and useful: matched GPT-4o mini comparisons against GEMBA-MQM and EAPrompt, official templates, code and prompts in the appendix, and a clean ablation showing dimension decoupling matters. Second, the headline ranking over learned metrics is not as solid as the abstract suggests. The debate topic, strategy, and round count were selected by maximizing meta scores on the WMT23 ZH-EN test set, and the same ZH-EN result is included in the average meta score that drives the abstract claims. No significance or variance is reported, and the margins are small.\n\nWhat's new: the three-stage pipeline (dimension partition, per-dimension pro-con debate with consensus, final judge) is a new combination, and the finding that coupled off-the-shelf debating frameworks hurt (Table 5) is a useful negative result. The ablation showing Stage 1 contributes the most is the real message.\n\nSoft spots. The selection-on-test-set issue is load-bearing. Tables 6, 7 and Figure 3 pick Severity/Consensus/R=3 on ZH-EN, then the same set appears in the headline average. EN-DE results provide some transfer evidence, but the headline mixes a tuned ZH-EN score with untuned EN-DE. Second, the debate component contributes almost nothing: removing Stage 2 changes ZH-EN meta by -0.006. The paper's title and framing emphasize multi-agent debate, but the evidence says dimension-splitting is the workhorse. Third, the HE-EN claim is wrong: Table 9 shows M-MAD's meta (0.787) is below MetricX-23 (0.807), XCOMET-QE-Ensemble (0.789) and MaTESe (0.792), yet the appendix says M-MAD is 'superior among all metrics'. That overstatement deserves a correction. Fourth, no standard errors, and the temperature=0 determinism doesn't quantify the test-set/configuration interaction. Also the case study in Table 10 shows M-MAD giving -2 vs gold -56; gold may be wrong, but the discrepancy is large and unresolved.\n\nBottom line: the paper deserves a serious referee. The selection issue is addressable with a held-out split or nested validation, and the debate-stage ablation should temper the framing. I'd send it out, with a request to fix the HE-EN claim and add uncertainty estimates. It's a solid empirical contribution for MT evaluation and LLM-as-a-judge communities, but the abstract oversells the debate mechanism.","headline":"Useful matched-baseline study of dimension-split LLM MT evaluation, but the headline ranking is weakened by test-set configuration selection and a debate component that contributes almost nothing.","tokens_in":22510,"tokens_out":3342,"would_cite":true,"duration_ms":29988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"M-MAD claims that an LLM-as-a-judge evaluator for machine translation can reach the accuracy of trained automatic metrics by splitting the MQM rubric into separate dimensions, having agents debate each dimension, and combining the debates…","keywords":["machine translation evaluation","LLM-as-a-judge","multi-agent debate","MQM","quality estimation","segment-level evaluation","WMT23","GPT-4o mini"],"falsifier":"Evaluate M-MAD with the exact ZH-EN-tuned settings on the WMT24 Metrics Shared Task (or any held-out set with MQM labels). If its meta-score no longer exceeds GEMBA-MQM, or falls more than one point below the XCOMET-QE ensemble, the reported margin is configuration-specific. A second, cheaper check: re-run the ZH-EN experiments with the debate topic fixed to \"Category\" instead of \"Severity\" and R=2; the paper's Table 6 predicts a drop of more than 0.14 in meta-score, so observing a smaller drop or a gain would refute the severity-debate mechanism.","tokens_in":21472,"feed_emoji":"🗣️","tokens_out":5983,"duration_ms":56584,"temperature":0.7,"pith_summary":"The paper proposes M-MAD, a three-stage framework that turns an LLM into a machine-translation evaluator. Instead of asking one model to apply the full MQM error rubric at once, it splits the rubric into four dimensions (accuracy, fluency, terminology, style), has pairs of agents debate the severity of errors in each dimension, and then has a judge merge the viewpoints into a final error list and MQM score. The claim is that this division and debate lets a cheap model like GPT-4o mini beat other LLM-as-a-judge methods and match trained reference-based metrics such as MetricX-23 and XCOMET on the WMT23 benchmark. If true, it would mean high-quality MT evaluation no longer requires training dedicated metric models or human-level annotation budgets, just orchestrated prompting.","feed_headline":"Debate turns LLM judges into top-tier MT metrics","feed_subtitle":"A three-stage, multi-agent framework with GPT-4o mini rivals trained metrics like MetricX-23 and XCOMET on WMT23.","key_machinery":"The load-bearing object is the three-stage M-MAD pipeline plus the MQM severity-weighted scoring formula. Stage 1 partitions the MQM rubric into $d=4$ dimensions (accuracy, fluency, style, terminology) with an independent few-shot evaluation agent per dimension. Stage 2 runs a two-agent pro-con debate on the severity of each detected error, with a consensus checker ending the debate early and a default rule that unresolved disagreements keep the supportive side's initial evaluation. Stage 3's judge agent merges dimension viewpoints, removes overlapping spans, and produces the final annotation set, from which the score is computed as $MQM_{score} = -5\\,n_{major} - 1\\,n_{minor}$. The mechanism that carries the argument is that decoupling removes the coupled-template bias of GEMBA-MQM-style prompts, and the severity-focused consensus debate corrects the over-severity bias that single-agent LLM judges show, as evidenced by the error-span prediction F1 (0.54 vs 0.37 for GEMBA-MQM) and the MQM score distribution that matches gold annotations.","core_discovery":"The central discovery is a \"neural network in natural language form\" — the paper's own analogy — in which the stages act as layers, agents as neurons, and their message exchanges as hidden states. The three stages are: first, decouple the MQM annotation guideline into four independent evaluation dimensions and have an initial agent produce fine-grained error annotations per dimension; second, run a two-agent pro-con debate on error severity within each dimension until consensus or three rounds, with a default toward minor severity; third, have a judge agent merge the dimension viewpoints, removing duplicate or overlapping error spans, and compute the final score with the weighted MQM formula ($w_{major}=5$, $w_{minor}=1$). The paper reports that this pipeline raises segment-level agreement with human judgments substantially over GEMBA-MQM and EAPrompt, and the overall meta-score on WMT23 ZH-EN reaches 0.808, ahead of every compared LLM-as-a-judge baseline and most learned metrics, trailing only the XCOMET-Ensemble among reference-based systems.","pith_inferences":["The configuration was chosen on the WMT23 ZH-EN test set; a natural next check is whether the same settings (severity topic, consensus, R=3) win on WMT24, which would rule out selection-on-test-set as the source of the gains.","The dimension-decoupling idea is general: any MQM-style or rubric-based evaluation (e.g., summarization, dialogue) could be split per criterion with per-criterion debate, and the paper's evidence suggests the gains come from the decoupling itself, not from translation-specific prompts.","The paper's own limitation note implies a testable extension: heterogeneous debating groups mixing strong closed models and weak open models might push performance higher than the homogeneous GPT-4o mini groups used here, and would also test whether debate gains scale with total reasoning budget.","The high token cost of repeated debates suggests a practical threshold: for deployment, one could measure the cost-quality tradeoff of reducing rounds to 2 or using a cheaper debater in Stage 1 with the expensive model only in the debate stage."],"forward_implications":["Using the same pipeline, an LLM-as-a-judge evaluator beats all other prompting-only evaluators and several trained metrics without any training data, on three language pairs of WMT23.","Because M-MAD is reference-free and training-free, MT evaluation could be run on demand for any language pair an LLM can translate, without collecting human-annotated metric training sets.","The configuration findings — severity debates beat category debates, consensus beats review-style strategies, three rounds is optimal — give concrete design rules for multi-agent evaluators beyond MT.","The decoupling insight implies that any coupled MQM-style prompt inherits a detectable error-severity bias, which explains the gap between system-level and segment-level performance in prior LLM-as-a-judge methods.","If the improvements hold on fresh test sets, the main obstacle to LLM-as-a-judge metrics surpassing learned metrics becomes cost, not capability."],"supporting_citations":[{"why":"Supplies the principal LLM-as-a-judge baseline, GEMBA-MQM, and the coupled MQM template that M-MAD decouples.","marker":"[Kocmi and Federmann, 2023a]"},{"why":"Provides the second baseline, EAPrompt, whose error-severity prompting M-MAD inherits and improves upon.","marker":"[Lu et al., 2024]"},{"why":"Supplies the WMT23 Metrics Shared Task datasets, human MQM ground truth, and the four-scenario meta-evaluation protocol that all reported scores come from.","marker":"[Freitag et al., 2023]"},{"why":"Defines the MQM annotation scheme and the severity-weighted scoring formula used in Stage 3.","marker":"[Freitag et al., 2021]"},{"why":"Supplies XCOMET, the trained reference-based metric that M-MAD aims to match; its ensemble is the only system scoring higher on the averaged meta-score.","marker":"[Guerreiro et al., 2024]"},{"why":"Supplies MetricX-23, the trained reference-based metric that M-MAD ties or slightly beats on WMT23.","marker":"[Juraska et al., 2023]"},{"why":"Provides the debating-with-persuasive-LLMs strategy whose direct adaptation to coupled MQM results underperforms, motivating M-MAD's dimension partitioning.","marker":"[Khan et al., 2024]"},{"why":"Offers the multi-agent debate framework whose direct application to coupled MQM results underperforms, motivating M-MAD's dimension partitioning.","marker":"[Chan et al., 2023]"}],"fun_headline_variants":["Debating agents turn LLM judges into top MT metrics","Multi-agent debate lifts LLM judges to rival learned metrics","M-MAD: Agent debates boost LLM judges to state-of-the-art MT eval","LLM judges match learned MT metrics via multi-agent debate","Agent-based debate pushes LLM judges past GEMBA and EAPrompt"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central premise is that the framework configuration tuned to perform best on the WMT23 ZH-EN test set — debating error severity, using consensus, with three rounds — also works for EN-DE and HE-EN and for future test sets; if that choice is noise, the reported improvements over baselines may not persist.","fun_headline_variants_meta":{"raw":{"variants":["Debating agents turn LLM judges into top MT metrics","Multi-agent debate lifts LLM judges to rival learned metrics","M-MAD: Agent debates boost LLM judges to state-of-the-art MT eval","LLM judges match learned MT metrics via multi-agent debate","Agent-based debate pushes LLM judges past GEMBA and EAPrompt"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1449,"prompt_tokens":997,"completion_tokens":452,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":361}},"tokens_in":613,"tokens_out":452,"duration_ms":5214,"temperature":1.0,"reasoning_tokens":361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:30:48.363497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate M-MAD with the exact ZH-EN-tuned settings on the WMT24 Metrics Shared Task (or any held-out set with MQM labels). If its meta-score no longer exceeds GEMBA-MQM, or falls more than one point below the XCOMET-QE ensemble, the reported margin is configuration-specific. A second, cheaper check: re-run the ZH-EN experiments with the debate topic fixed to \"Category\" instead of \"Severity\" and R=2; the paper's Table 6 predicts a drop of more than 0.14 in meta-score, so observing a smaller drop or a gain would refute the severity-debate mechanism.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the debating-with-persuasive-LLMs strategy whose direct adaptation to coupled MQM results underperforms, motivating M-MAD's dimension partitioning."}],"review_version":1}