{"id":"52a86ca5-319e-458c-a597-dd42fe0cccb1","arxiv_id":"1909.01701","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new multilingual audit-domain test suite shows that news-trained MT systems appear fluent but fail on domain-specific terms and document-level semantics, including party identity in agreements.","lead":"Audit reports from the Czech Supreme Audit Office were turned into a machine translation test suite and used in WMT19. The evaluation finds that news-trained MT systems translate general audit text well, but lose critical details such as the identity of the parties in agreements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sublease claim may be an artifact of the source: the English 'tenant/lessee' distinction is never validated against the original Czech, so MT's term collapse could be faithful to an ambiguous input.","rationale":"The reader identified the single-annotator, no-IAA gold standard and the single-document scope as the weakest premises. I agree those are real limitations, and the authors themselves flag them in the text ('we did not collect enough annotations to reliably measure it'; 'small sample size'). But the more load-bearing premise, in my reading, is the validity of the source-side markables. The paper's own Section 5 states that the English sublease agreement is a non-professional translation from Czech. The entire 'identity of parties' result depends on treating the English words 'tenant' and 'lessee' as a reliable encoding of two distinct legal roles. If the English source is lexically conflated—a plausible outcome when translating Czech 'nájemce' into English—then MT systems reproducing that conflation are not failing to preserve the agreement's semantics; they are preserving the source's ambiguity. The authors do not report any check of this. This is not a question of annotator subjectivity; the Clash label in Table 16 may be perfectly reliable relative to the source, but the source itself may be the wrong authority. The concrete test I propose settles this by aligning the English source with the original Czech, which is available in the released suite and was used in the evaluation. If the test shows a clean, consistent tenant/lessee distinction, the abstract's claim survives for this case study and the paper remains a useful resource—still best read as a case study, not a universal law, and still CONDITIONAL because of the single-document and single-annotator basis. If the test shows conflation, the abstract should be softened to say the systems preserve the source's ambiguity, or should restrict the claim to cases where the source is unambiguous. Either way the reader's CONDITIONAL verdict stays; I do not see grounds for REJECT, since the test suite release and the manual protocol are independently valuable.","tokens_in":16493,"tokens_out":10513,"duration_ms":105748,"concrete_test":"Use the released suite at https://github.com/ELITR/wmt19-elitr-testsuite. Extract every occurrence of 'tenant' and 'lessee' in the English sublease source and align each occurrence to the corresponding clause or term in the original Czech text, which the authors state exists and used in evaluation. Determine whether each English term maps consistently to a distinct party or role and never to the same Czech expression. If the source conflates the terms or maps them inconsistently, re-annotate Table 16 with ambiguous source mentions excluded or resolved from the Czech original; if the clash counts for MT systems fall materially, the headline 'completely fail' claim does not survive. Report the result as an erratum or as a conditioning note in the paper.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest abstract claim—that MT systems 'completely fail in preserving the semantics of the agreement, namely the identity of the parties'—rests on the tenant/lessee clash counts in Table 16. Section 5 reveals that the English source of the sublease agreement is itself 'a (non-professional) translation from Czech.' The paper reports no verification that this English source consistently uses 'tenant' for one contracting party and 'lessee' for the other. If the non-professional translation uses the two terms interchangeably (both plausibly rendering Czech 'nájemce'), then an MT output that translates both as 'nájemce' is semantically faithful to the source, and the observed 'clash' is inherited from the input rather than introduced by the MT systems. The central causal claim would then be unsupported; only a claim about ambiguity preservation would remain. This is distinct from, and prior to, the inter-annotator reliability concern: even a perfectly reliable annotator cannot rescue a markable set whose source-side referents were never checked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the SAO WMT19 test suite, a publicly released set of multilingual audit-report documents used as a WMT19 test suite for Czech, English, and German translation, and reports automatic and manual evaluations of the participating MT systems. The automatic scores show reasonable domain-level performance, while a manual evaluation with domain experts indicates that subtle terminological errors are missed by reference-based automatic metrics. In a separate analysis of one sublease agreement, the authors report that all tested systems translate both \"tenant\" and \"lessee\" as Czech \"nájemce,\" collapsing the distinction between the contracting parties, and conclude that even the best systems completely fail to preserve the semantics of the agreement regarding party identity.","tokens_in":16640,"tokens_out":8086,"duration_ms":81568,"significance":"The paper's main contribution is a useful public test suite for a non-news domain, together with a detailed manual annotation protocol and a rare attempt to evaluate document-level semantic consistency in MT. The authors are commendably explicit about limitations: they report small sample sizes, standard deviations on automatic scores, the absence of reliable inter-annotator agreement, and the use of a single annotator for some language pairs. If the party-identity claim can be substantiated with a proper source-side validation, the work provides a concrete, falsifiable case that strong sentence-level NMT systems can destroy legally crucial role distinctions while scoring well on generic metrics. The public release of the test suite and annotation materials is a valuable resource for the community.","major_comments":[{"comment":"The central empirical claim that all systems conflate the two contracting parties by translating both \"tenant\" and \"lessee\" as \"nájemce\" presupposes that the English source consistently uses \"tenant\" for one party and \"lessee\" for the other. Section 5 states that the English source \"was in fact a (non-professional) translation from Czech,\" but the paper never checks the Czech original for how the parties are referred to and whether the English terms align with distinct Czech referents. If the non-professional English translation uses the two terms interchangeably, then an MT output that maps both to \"nájemce\" may be faithful to the source, and the observed clash would be inherited rather than introduced by MT. Please add a term concordance of the Czech original, the English source, and the occurrences counted in Table 16, and verify that every \"clash\" corresponds to a genuine role distinction in the source. If the check does not support the premise, the abstract and Section 7 claims must be weakened to a claim about preserving source ambiguity rather than party identity.","section":"Section 5, Table 16"},{"comment":"The manual evaluation underlying the paper's stronger conclusions has very limited reliability evidence, which the authors acknowledge in Section 4.1 (\"we did not collect enough annotations to reliably measure it\"). Only three segments were double-annotated, with known disagreements, and all English-to-German and German-to-English evaluations were performed by a single annotator. For the sublease agreement, the annotation was described as \"partially blind\" and performed by the \"main annotator\" without a second annotation. This does not invalidate the findings, but it does mean that the abstract statement that automatic MT evaluation with one reference is \"practically useless\" is stronger than what the reliability evidence supports. I recommend reporting per-criterion agreement on the double-annotated segments, providing a second annotation of the party-reference marking in the sublease agreement, or explicitly restricting the automatic-evaluation claim to this test suite and to the specific error types studied rather than presenting it as a general conclusion.","section":"Section 4.4 and Section 5.4"},{"comment":"The paper argues that automatic evaluation with one reference is \"practically useless\" for detecting the fine-grained semantic errors it describes, but it never reports automatic scores for the sublease agreement, which is the document where the most severe error occurs. Without a direct comparison of automatic metric scores on that document against the manual error counts, the reader cannot assess whether the metrics fail specifically where the semantic distinction matters. I suggest either including the relevant automatic scores for the sublease agreement in Section 5, or adding a numerical comparison (e.g., correlation between automatic scores and expert judgments across the audit documents) to substantiate the general claim; otherwise the wording should be softened to \"automatic metrics did not flag these errors in our evaluation.\"","section":"Section 4.2 and Section 7"}],"minor_comments":[{"comment":"The sentence referring to \"the main Findings of WMT19 paper\" has no corresponding reference in the bibliography; please add the WMT19 Findings paper (Barrault et al., 2019).","section":"Section 4.4.1"},{"comment":"The sentence \"The original Czech text was evaluated with all other WMT19 systems as if it was one of the systems\" is unclear, since the original Czech text cannot be an output of an English-to-Czech system; please clarify whether it was used as a reference, a control condition, or something else.","section":"Section 5.1"},{"comment":"The header \"Apartement in Question\" contains a spelling error; it should be \"Apartment in Question.\"","section":"Table 15"},{"comment":"There is a typo in \"we demanded the this particular choice\"; it should read \"we demanded this particular choice.\"","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful resource paper from the WMT19 test-suite side: it contributes a new multilingual audit-domain test set, releases it, and reports a careful manual evaluation that finds a concrete, reproducible failure mode. The new thing is the SAO Test Suite plus a markable-based annotation protocol for agreement documents. The central empirical finding is that on a sublease agreement, every tested system translates both 'tenant' and 'lessee' as the same Czech word 'nájemce', making the two contracting parties indistinguishable. The paper does this work honestly: automatic scores come with standard deviations, the authors flag the lack of a full inter-annotator agreement study, they disclose that the online systems might have seen the test suite, and they are explicit that the sublease result comes from a single sample document.\n\nThe main soft spot is the one the stress-test note points at. The English source for the sublease agreement is itself a non-professional translation from Czech, and the paper never explicitly verifies that the source uses 'tenant' and 'lessee' for two genuinely distinct parties. That check matters because if the source used the terms interchangeably, the MT output collapsing them to 'nájemce' could be faithful. On the evidence in the paper I don't think that is the case: the Reference row in Table 16—which is the original Czech text evaluated as a candidate—gets 16 of 17 party references right, and the authors explain the legal distinction between nájemce and podnájemce in subleases. So the source probably did distinguish the parties. But the paper should say so explicitly and show the source-side markable assignment; right now a reader can't fully rule out an inherited ambiguity.\n\nTwo smaller issues. The agreement evaluation appears to have one main annotator, and the paper doesn't state how many people did the sublease annotations or whether any consistency check was done; for a load-bearing qualitative claim, that's worth adding. And the claim that automatic MT evaluation is 'practically useless' is a reasonable qualitative takeaway but only weakly evidenced; one reference is clearly insufficient, but the paper doesn't test metric behavior in a way that would support a stronger statement.\n\nOverall this is a solid resource paper with a real warning for MT evaluation practice. It deserves a serious referee; the needed revisions are small, mostly adding the source-side validation and clarifying the annotation setup. I'd send it to review rather than desk reject.","headline":"Useful test suite, honest manual eval, and a party-identity finding that is probably real but needs an explicit source-side check.","tokens_in":17197,"tokens_out":6243,"would_cite":true,"duration_ms":58036,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current machine translation systems, even the best in the 2019 news shared task, fail to preserve the identity of contracting parties in a sublease agreement; single-reference automatic metrics are practically useless for detecting this.","keywords":["machine translation","audit reports","test suite","domain adaptation","manual evaluation","terminology consistency","contracting parties","sublease agreement"],"falsifier":"Translate a set of originally-English sublease agreements into Czech with these systems and check every mention of 'tenant' and 'lessee': if any system consistently uses two distinct terms (for example, 'nájemce' and 'podnájemce') for the two roles, the claim that even the best systems completely fail at party identity is refuted.","tokens_in":16304,"feed_emoji":"🤝","tokens_out":9567,"duration_ms":90869,"temperature":0.7,"pith_summary":"The paper builds a public test suite of Czech, English, and German audit reports and runs WMT19 news-trained machine translation systems on it. Its central claim is that, on ordinary audit-report text, these systems look near-perfect to a non-expert and score well automatically, but expert auditors find fine-grained terminological errors that a single reference translation cannot reveal. On a further sample document, a sublease agreement, the paper reports that every tested system translated the two contracting parties 'tenant' and 'lessee' as the same Czech word, destroying the meaning of the agreement. The upshot would matter because it locates a concrete failure mode—preserving the semantic identity of legal roles across a document—that current architectures do not address and current evaluation cannot see.","feed_headline":"Every MT system tested merged tenant and lessee into one word","feed_subtitle":"None of the systems kept the two parties distinct; automatic scores with one reference missed it.","key_machinery":"The load-bearing instrument is a manual, markable-based annotation protocol applied to one sublease agreement. The paper fixes a list of markables—named entities, dates, numbers, and document-specific legal terms such as 'tenant', 'lessee', 'supplement', 'equipment', and 'amenities'—and for each machine translation decides whether each occurrence is correct, wrong, missing, or clashing with another term. The decisive check is the term clash: 'tenant' and 'lessee' must map to two distinct Czech legal words ('nájemce' versus 'podnájemce') whenever both roles appear, otherwise a reader cannot tell which contracting party is which. This protocol sits on top of a trilingual parallel test suite of audit reports that the paper cleaned and released for reuse.","core_discovery":"On its own terms, the paper's discovery is that the ceiling for current MT is not where sentence-level fluency suggests. In the sublease document, every one of the evaluated systems used a single Czech word, 'nájemce', for both 'tenant' and 'lessee', even though Czech has distinct terms ('nájemce' versus 'podnájemce') for the two roles in a sublease. Because a contract's meaning depends on knowing which party is which, all translations became effectively incomprehensible at the one point that matters most. The paper further shows that automatic scores, all computed against one reference translation, cannot flag this: the reference itself sometimes uses acceptable variants, and the clash is a consistency and factual error rather than an n-gram difference. Manual evaluation by domain professionals was required even to see the failure.","pith_inferences":["The same party-identity failure should be expected in any target language that has distinct terms for landlord-tenant versus tenant-subtenant relationships; building a German version of the sublease test would be a direct check.","A quantitative extension would be to define an 'entity-identity preservation' score from the markable annotations—the percentage of role mentions translated to the correct distinct term—and test whether existing reference-free semantic metrics correlate with it.","Because the English source was itself translated from Czech, the reference translation carried a few errors; a cleaner test using an originally-English agreement with a vetted legal translation might make the systems' failure even sharper."],"forward_implications":["High automatic scores on domain text should not be read as evidence that a system understands the domain; the audit-report translations scored near the top while hiding serious terminology errors.","A contract-style document with two roles that must remain distinct is a cheap, high-signal probe for evaluating MT, since party identity is a semantic property that sentence-level metrics miss.","Terminology consistency across a document is not enforced by current systems; translating the same source term differently, or two source terms identically, can invert the meaning of a legal text.","For practical deployment in legal or audit settings, machine translation needs expert-in-the-loop correction or terminology lists, because neither automatic metrics nor unassisted MT output can guarantee role identity."],"supporting_citations":[{"why":"Reports that the best WMT18 system significantly outperformed humans at the sentence level, establishing the pedigree of the systems that then fail on the sublease agreement.","marker":"Bojar et al., 2018"},{"why":"Describes the document-level Transformer systems used in the evaluation and their multi-sentence training, the closest existing approach to document-level consistency.","marker":"Popel et al., 2019"},{"why":"Defines BLEU, the primary single-reference automatic metric whose insensitivity to terminological clashes underlies the claim that automatic evaluation is practically useless.","marker":"Papineni et al., 2002"},{"why":"Defines TER, another one-reference edit-rate metric in the automatic evaluation tables; it too cannot see the party-identity clash.","marker":"Snover et al., 2006"},{"why":"Provides the hunalign sentence-alignment tool used to build the trilingual audit-report test suite.","marker":"Varga et al., 2005"},{"why":"Supplies the domain-invariance hypothesis that motivates testing news-trained systems on audit reports.","marker":"Narayanan et al., 2018"}],"fun_headline_variants":["All MT systems conflate tenant and lessee in sublease","MT fails to distinguish tenant vs lessee in contracts","Reference-based scores blind to MT's legal errors","MT looks fluent but loses tenant-lessee identity","Manual review needed to catch MT's legal errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The universal-failure conclusion rests on one sublease document and a small set of domain-expert annotations, with no formal measure of annotator agreement and a single annotator for the German directions.","fun_headline_variants_meta":{"raw":{"variants":["All MT systems conflate tenant and lessee in sublease","MT fails to distinguish tenant vs lessee in contracts","Reference-based scores blind to MT's legal errors","MT looks fluent but loses tenant-lessee identity","Manual review needed to catch MT's legal errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001067,"raw_usage":{"total_tokens":4412,"prompt_tokens":829,"completion_tokens":3583,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":3508}},"tokens_in":445,"tokens_out":3583,"duration_ms":23722,"temperature":1.0,"reasoning_tokens":3508,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:09:02.742201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Translate a set of originally-English sublease agreements into Czech with these systems and check every mention of 'tenant' and 'lessee': if any system consistently uses two distinct terms (for example, 'nájemce' and 'podnájemce') for the two roles, the claim that even the best systems completely fail at party identity is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports that the best WMT18 system significantly outperformed humans at the sentence level, establishing the pedigree of the systems that then fail on the sublease agreement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the document-level Transformer systems used in the evaluation and their multi-sentence training, the closest existing approach to document-level consistency."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines BLEU, the primary single-reference automatic metric whose insensitivity to terminological clashes underlies the claim that automatic evaluation is practically useless."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines TER, another one-reference edit-rate metric in the automatic evaluation tables; it too cannot see the party-identity clash."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the hunalign sentence-alignment tool used to build the trilingual audit-report test suite."}],"review_version":1}