{"id":"dff1a986-16e5-420b-bcae-36230da352c2","arxiv_id":"2506.13468","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A position survey calling for human-centered machine translation, synthesizing translation studies and HCI to broaden MT evaluation and design beyond benchmark quality.","lead":"This paper argues that machine translation should be designed around the people who use it, not just around translation accuracy. It surveys research from translation studies and human-computer interaction to propose a research agenda for evaluating and building translation tools that fit real-world tasks and contexts.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim's impact rests on an unverified causal path: the paper's own cited interventions (Zouhar 2021; Mehandru 2023) improve confidence or reliance but not objective quality or error detection.","rationale":"The reader's ACCEPT verdict is defensible for a position paper: the normative claim that MT should be recontextualized through Translation Studies and HCI is well-supported by a broad literature review, and the Limitations section appropriately acknowledges scope constraints. The weakest point is the implicit causal claim that addressing socio-technical gaps will improve real-world value. The reader identified a closely related version (translation quality as the binding constraint); my check sharpens it using internal evidence from the paper itself: several cited interventions shift perceptions or trust without improving objective task outcomes. This does not make the paper internally inconsistent, because the authors explicitly call for more research, but it does mean the impact promise in Section 9 is not established by the cited evidence. Since the survey is valuable and should be published, I recommend CONDITIONAL acceptance rather than outright ACCEPT: the authors should explicitly frame the efficacy of human-centered interventions as an open empirical question and temper the claim that this approach 'promises greater real-world impact' until the mixed intervention results are reconciled.","tokens_in":26539,"tokens_out":4621,"duration_ms":51134,"concrete_test":"Code each intervention study cited in Sections 4 and 8 (e.g., Zouhar et al. 2021; Mehandru et al. 2023; Koehn 2010; Xu et al. 2014; Gao et al. 2015; Grissom et al. 2024) for whether it reports a significant improvement on an objective task-performance measure (comprehension accuracy, clinical error detection, post-edit quality, mutual understanding) versus only subjective confidence, trust, or preference. If a majority of coded studies show only subjective gains, the paper's headline promise should be relabeled as an open hypothesis rather than an expected outcome.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central argument is normative: MT research should broaden evaluation and design around situated use, user literacy, and risk communication. To deliver the promised 'greater real-world impact' (Section 9), human-centered interventions must actually improve real-world outcomes. That is the load-bearing step, and the paper's own evidence is mixed. Zouhar et al. (2021) find that backtranslation feedback raises user confidence but not the quality of the text produced; Mehandru et al. (2023) find that quality-estimation feedback improves physicians' reliance but fails to flag the most clinically severe errors. These results are presented honestly, but their implications are not reconciled: if the main measurable effect of a proposed intervention is miscalibrated confidence rather than better decisions, the design agenda does not yet have direct empirical support. Likewise, Asscher and Glikson (2021) show that labels shift perceived quality, which demonstrates malleability but not that accurate calibration can be achieved. The paper therefore establishes a plausible and well-organized research agenda, but the causal path from human-centered design and evaluation to improved outcomes is an assumption, not a finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that machine translation (MT) research should adopt a human-centered approach that broadens the goals of MT systems beyond producing fluent and adequate translations, toward helping users weigh risks and benefits and aligning system design with communicative goals. Drawing on Translation Studies and HCI, the paper surveys contexts of MT use, MT literacy, empirical findings on human-MT interaction, translation ethics, and then proposes research directions for more situated evaluation and for interaction and design, illustrated with a healthcare case study. The paper contains no new experiments; its contribution is a cross-disciplinary synthesis and a research agenda.","tokens_in":26714,"tokens_out":5262,"duration_ms":51573,"significance":"The paper's main strength is the breadth and timeliness of its synthesis: it brings together evidence from Translation Studies (e.g., MT literacy, translation briefs, ethics) and HCI (e.g., trust calibration, seams, mediated communication) that is often absent from MT benchmarking discussions. It carefully connects the surveyed literature to concrete design and evaluation proposals, such as evaluation briefs, augmented outputs, two-output interfaces, and risk-communication tools. The paper is honest about mixed empirical results, including Zouhar et al.'s finding that backtranslation feedback increases confidence but not quality and Mehandru et al.'s finding that quality estimation fails to flag the most clinically severe errors. If the field adopts this agenda, MT research would prioritize stakeholder-centered evaluation and interaction design alongside benchmark quality; this is a plausible and productive reorientation.","major_comments":[{"comment":"The conclusion states that a human-centered approach 'promises greater real-world impact' (Section 9), but the causal path from the proposed interventions to improved outcomes is not established by the cited evidence. In Section 4, Zouhar et al. (2021) show that backtranslation feedback increases user confidence without improving the quality of the produced text; in Section 8, Mehandru et al. (2023) find that quality-estimation feedback improves physicians' reliance but fails to detect the most clinically severe errors. These results do not yet show that the proposed human-centered features improve decisions or outcomes. The paper should explicitly identify the need to validate this causal chain as a first-order research question, and soften the 'promises' wording accordingly. This is a framing issue rather than a flaw in the survey, but it is load-bearing because the promised real-world impact is part of the central motivation.","section":"Sections 4, 8, 9"}],"minor_comments":[{"comment":"The claim that 99.97% of MT users are not professional translators (citing Nurminen, 2021a) should state the basis of this estimate, since it is a striking and widely-quoted number.","section":"Section 2"},{"comment":"'Subsequent phrases include usability testing' appears to be a typo for 'phases'; please correct.","section":"Section 6"},{"comment":"'Zouhar et al. (2021) studies' should be 'study' for subject-verb agreement.","section":"Section 4"},{"comment":"The entry for O'Brien, Simard, and Goulet (2018) appears twice with slightly different capitalization and page ranges; the duplicate should be merged.","section":"References"},{"comment":"Some diacritics are corrupted in author names (e.g., 'Skadi n, a' for Skadiņa, 'Vasi l.jevs' for Vasiljevs); the bibliography should be cleaned up.","section":"References"},{"comment":"The paragraph on 'compartmentalization of MT ethics from general AI ethics' (citing Asscher, 2025) introduces a potentially useful distinction, but the practical implications for MT design and evaluation are left implicit; a sentence connecting this to the following design sections would improve readability.","section":"Section 5"},{"comment":"The discussion of the 'translation brief' is clear, but the paper does not discuss how a brief might be operationalized in an MT interface; consider a pointer to Section 7's 'evaluation brief' concept.","section":"Section 3"}],"recommendation":"minor_revision","confidential_remarks":"This is a broad position paper from a large, well-known author group. The survey appears balanced and the evidence is cited honestly. If the journal prefers empirical or technical contributions, the fit should be considered; but as a roadmap for a research program, it is likely to be influential. The heavy citation of co-authored work is expected for a synthesis by leaders in the area; I did not see evidence of self-citation crowding out contradictory findings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll be direct: this is a solid, well-organized position survey, not a new empirical result. Its value is in the synthesis—it brings together MT, Translation Studies, and HCI in a way that genuinely clarifies why human-centered MT should look different from benchmark-driven MT. The sections on MT literacy, trust, post-editing, and ethics are well structured, and the healthcare case study is a good concrete anchor. I particularly credit the authors for stating their limitations plainly (no multimodality, not exhaustive) and for hedging recommendations as directions rather than certainties.\n\nThe soft spot is exactly where the stress-test note lands. The paper's conclusion promises 'greater real-world impact' from this reorientation. That's a causal claim—if we redesign evaluation and interaction around situated use, outcomes for real users will improve. The evidence cited for that path is mixed. Zouhar et al. (2021) show backtranslation feedback increases user confidence but not the quality of the produced text. Mehandru et al. (2023) show QE improves reliance but fails to catch the most clinically severe errors. The paper reports these honestly, but it doesn't reconcile them with the promised payoff. Maybe the point is that these are early steps, and that's fine for a research agenda. But the paper doesn't quite say that. It reads at times as if the design agenda is already validated. It also doesn't seriously weigh the alternative that translation quality itself is the binding constraint in many high-stakes settings.\n\nThat caveat aside, the paper is a useful contribution. It gives researchers a shared vocabulary and a concrete list of open problems. The citations are real and relevant; the self-citations are to empirical studies, so I don't see a circularity problem. I'd bring this to a reading group and probably cite it when writing about evaluation or user interaction.\n\nFor peer review: yes, send it out. A good referee should focus on the gap between agenda and evidence—asking the authors to state explicitly which recommendations are hypotheses, and how they would test them. That's a fixable issue, not a fatal one.","headline":"A solid, well-organized position survey that makes a real case for human-centered MT, but its promised impact rests on an assumption that its own mixed evidence does not yet support.","tokens_in":27370,"tokens_out":2179,"would_cite":true,"duration_ms":22144,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that machine translation should be treated as a human-centered, socio-technical process, evaluated for fitness-for-purpose in context and designed to help users weigh risks and benefits, rather than optimized for…","keywords":["machine translation","human-centered AI","translation studies","human-computer interaction","MT literacy","MT evaluation","trust calibration","socio-technical systems"],"falsifier":"A controlled field study in a high-stakes setting would settle the matter: if a human-centered MT system with quality-estimation feedback, backtranslation, and side-by-side source text does not reduce clinically harmful errors or improve appropriate trust relative to a standard generic MT tool, the paper's central claim would be undermined.","tokens_in":26318,"feed_emoji":"🌐","tokens_out":9722,"duration_ms":82053,"temperature":0.7,"pith_summary":"Machine translation means translating automatically, but this paper argues that the field's real-world value is held back by a socio-technical gap: most users are not professional translators, they often lack the knowledge to judge whether a translation is reliable, and evaluations measure generic quality rather than whether the output works for the reader. The authors propose a human-centered approach that imports concepts and methods from Translation Studies and Human-Computer Interaction, treating MT as a situated, interactive process aligned with users' communicative goals. If the argument is right, MT evaluation should shift from benchmark scores to fitness-for-purpose and stakeholder impact, and MT tools should help lay users calibrate trust, manage risk, and interact with translations rather than simply producing fluent text.","feed_headline":"Machine translation should center people, not just quality","feed_subtitle":"A cross-disciplinary survey argues MT evaluation and design must match how lay users weigh risks and use translations","key_machinery":"The load-bearing concept is the 'translation brief', a professional translator's specification of why, for whom, and how a translation will be used, which the paper generalizes into an 'evaluation brief' for MT and into design requirements for context-aware systems. This concept, together with MT literacy and human-centered design methods from HCI, converts MT from one-shot sequence transduction into an iterative, socio-technical process whose success is measured by whether it supports informed, contextually appropriate communication.","core_discovery":"The central claim is that building genuinely useful MT systems requires recontextualizing evaluation and design around human users and their contexts, not just around translation quality. Drawing on empirical work, the paper shows that users cope with imperfect MT through strategies like backtranslation, simplified self-expression, and holistic conversation understanding, but these strategies carry costs; trust is shaped by labels and interface cues as much as by accuracy; and errors that read smoothly can mislead more than obvious ones. It therefore argues that MT research should adopt the 'translation brief'—specifying purpose, audience, and use—as a model for evaluation, and should design systems that support MT literacy, risk management, and iterative interaction, with human studies in real tasks guiding the process.","pith_inferences":["Editorial inference: if this framing is adopted, common quality benchmarks such as BLEU cannot be treated as proxies for real-world value; a testable extension would compare how well benchmark rankings predict user-centered outcomes such as appropriate trust, task success, and harm across domains.","Editorial inference: the healthcare case study points toward a regulatory conclusion the authors do not state—high-stakes MT should be held to fitness-for-purpose standards backed by vetted phrase scaffolds, verifiable outputs, and interaction design, rather than to general-purpose quality scores.","Editorial inference: as translation becomes embedded in general-purpose LLM workflows, often covertly, the paper's logic implies that disclosure and transparency—telling users when and how MT was used—become load-bearing design features on par with translation quality."],"forward_implications":["MT evaluation should shift from generic benchmark scores toward situated assessments of fitness-for-purpose and stakeholder impact, using 'evaluation briefs' that specify purpose, audience, and intended use.","MT systems for lay users should embed supports that calibrate trust, such as quality estimation, backtranslation, multiple outputs, and side-by-side source text, because evidence shows users struggle to assess reliability.","Design should treat MT as an iterative, interactive process—pre-editing, post-editing, and user control—rather than a one-shot sequence-to-sequence output.","Research should include needs-finding, co-design, and human studies in context, measuring interpersonal and communicative outcomes as well as task performance.","MT literacy should be promoted as a design goal, since most users are not professional translators and many hold misconceptions about translation."],"supporting_citations":[{"why":"Supplies the socio-technical gap concept used to frame the divide between MT development and real-world use.","marker":"(Ackerman, 2000)"},{"why":"Provides the assimilation/dissemination/communication taxonomy of MT use that structures the paper's context analysis.","marker":"(Hovy et al., 2002)"},{"why":"Establishes machine translation literacy as a pressing need for non-professional users.","marker":"(Bowker and Ciro, 2019b)"},{"why":"Defines MT literacy as knowing how MT works, when it is useful, and what using it implies.","marker":"(O'Brien and Ehrensberger-Dow, 2020)"},{"why":"Documents that about 99.97% of MT users are not professional translators, anchoring the focus on lay users.","marker":"(Nurminen, 2021a)"},{"why":"Shows fluency errors lower user trust more than adequacy errors, motivating risk-calibration design.","marker":"(Martindale and Carpuat, 2018)"},{"why":"Demonstrates that backtranslation feedback boosts user confidence without improving output quality, supporting the need for reliability signaling.","marker":"(Zouhar et al., 2021)"},{"why":"Shows quality estimation improves physicians' reliance on MT but misses the most clinically severe errors, grounding the healthcare case study.","marker":"(Mehandru et al., 2023)"},{"why":"Documents clinically harmful errors in Google Translate discharge instructions, motivating the clinical reliability agenda.","marker":"(Khoong et al., 2019)"}],"fun_headline_variants":["MT users cope with errors — design for that, not just quality","Translation brief: The missing metric for MT evaluation","Why MT evaluation should look beyond accuracy","Human-centered MT: purpose, audience, and context matter","Lay users backtranslate and simplify — design MT for that"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the socio-technical gaps—low MT literacy, miscalibrated trust, and context-insensitive evaluation—are the main barriers to MT's real-world value and can be narrowed by redesigning systems and evaluation; if translation quality itself were the binding constraint, the proposed redesign would not deliver the promised gains.","fun_headline_variants_meta":{"raw":{"variants":["MT users cope with errors — design for that, not just quality","Translation brief: The missing metric for MT evaluation","Why MT evaluation should look beyond accuracy","Human-centered MT: purpose, audience, and context matter","Lay users backtranslate and simplify — design MT for that"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1461,"prompt_tokens":766,"completion_tokens":695,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":382,"completion_tokens_details":{"reasoning_tokens":617}},"tokens_in":382,"tokens_out":695,"duration_ms":6010,"temperature":1.0,"reasoning_tokens":617,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:59:50.494123+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled field study in a high-stakes setting would settle the matter: if a human-centered MT system with quality-estimation feedback, backtranslation, and side-by-side source text does not reduce clinically harmful errors or improve appropriate trust relative to a standard generic MT tool, the paper's central claim would be undermined.","supporting_citations":[],"review_version":2}