{"id":"aa0c10b5-6991-4021-90fe-a5e31872dfbe","arxiv_id":"2412.08392","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Multilingual LM evaluations that mix English prompts with target-language text measure task performance boosted by an English interface, not true multilingual language understanding.","lead":"This position paper distinguishes two roles English plays in multilingual language model evaluation: as an interface to boost task performance and as a natural language for testing understanding. It argues that mixed English and target-language prompts are imprecise and recommends native-language or natural code-switched evaluation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'inherently imprecise' conclusion is conditional on an unstated deployment goal; if the evaluation target is task performance under an English-prompt interface, mixed prompts are precise, so the blanket recommendation does not follow.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the recommendation depends on prioritizing language understanding over task performance, and the paper never refutes the task-performance perspective. My stress-test sharpens this into a precise internal gap: Section 3 claims to show imprecision 'even if maximizing task performance is the goal,' but the arguments given establish imprecision only relative to MLU. If task performance is the goal and the deployment interface is the mixed prompt itself, the evaluation is not imprecise at all. The paper's normative preference for MLU is stated clearly and is defensible as a position, but the blanket wording of the conclusion overreaches that position. The paper remains a valuable position paper, and the reader's CONDITIONAL verdict is appropriate: the central claim should be qualified to make the deployment condition and evaluative goal explicit. No change to the reader's verdict is needed.","tokens_in":8383,"tokens_out":3063,"duration_ms":33517,"concrete_test":"Fix a deployment condition. Take SIB-200 with 205 languages and a single multilingual LM. Define the target construct T as task performance under the exact prompt format the system will use in production. Compare (a) mixed English-target prompt and (b) native target-language prompt as estimators of T. If the production format is the English mixed prompt, then (a) is unbiased by construction, contradicting the claim that it is 'inherently imprecise' for task performance. The paper needs to either state the production interface or restrict the claim to MLU. A minimal analytical check: re-read Section 3's examples and replace 'task performance' with 'performance under a target-language-only deployment'; the examples no longer support the conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is in Section 3 and the Conclusion: the authors claim a mixed-prompt is 'inherently imprecise ... even if maximizing task performance is the goal.' But throughout Section 3, imprecision is established relative to multilingual language understanding (MLU), e.g., 'we arguably do not test MLU' and 'more than just the task is evaluated.' For an evaluator whose target construct is task performance in a system that will actually be deployed with English as the prompt language, the mixed prompt is not a confound; it is the operational environment. The paper grants in Section 2 that task performance is a widespread, practical perspective, but never refutes it; it only calls it 'seemingly at odds' with the NLU label. The conclusion that mixed prompts yield 'imprecise or misleading evaluations' therefore relies on an unargued normative priority: language understanding is the correct object of evaluation. Without that priority, the recommendation to abandon mixed prompts does not follow. This is not merely a disagreement with consensus; it is an internal gap between the stated goal (task performance) and the metric used to judge imprecision (MLU).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that English used in multilingual language-model evaluation can play two distinct roles: as an interface (a prompt format chosen to maximize task performance) and as a natural language (a target of linguistic understanding). The authors contend that mixing an English instruction frame with a target-language passage, which they call a mixed-prompt, is unnatural rather than genuine code-switching, and that it evaluates more than the intended task or than multilingual language understanding (MLU). They illustrate this with examples from SIB-200, IrokoBench, and a machine-translation prompt from Hendy et al. They conclude that mixed-prompt evaluations are imprecise or misleading, even when the stated goal is task performance, and recommend moving toward native target-language prompts or natural code-switched prompts.","tokens_in":8729,"tokens_out":7675,"duration_ms":83709,"significance":"If the conceptual distinction were fully sustained, the paper would make a useful contribution by giving researchers a vocabulary for separating interface-level task performance from language understanding in multilingual benchmarks. Its strengths are the clear two-role taxonomy, the concrete worked examples, the explicit acknowledgment of the task-performance perspective, and the transparency about nonstandard terminology. The paper makes no empirical claims and ships no code, which is appropriate for a position paper. The central weakness is that the argument's strongest conclusion—that mixed prompts are imprecise even for task performance—is not supported by the definitions provided; the analysis establishes imprecision only relative to MLU, not relative to a task-performance construct. If that overreach is corrected, the paper can serve as a helpful caution against conflating benchmark scores with language understanding.","major_comments":[{"comment":"The paper's strongest conclusion, that mixed prompts yield 'imprecise or misleading evaluations, even if the ultimate goal was to evaluate and improve task performance,' is not actually established by the preceding argument. In §3, the imprecision is defined relative to multilingual language understanding ('we arguably do not test MLU'), while the task-performance perspective in §2 is acknowledged as practical but never given a target definition. If the evaluation construct is task performance in a system that is deployed with an English instruction template, a mixed-prompt evaluation measures exactly the behavior of interest; the English frame is the operational environment, not an extra variable. The authors need to either reject the task-performance perspective on substantive grounds or qualify the conclusion so that 'imprecise or misleading' applies only to evaluations that claim to measure MLU. As written, the recommendation to abandon mixed prompts overreaches its premises.","section":"§3, first paragraph; Conclusion"},{"comment":"The bullet list stating that the IrokoBench prompt tests 'code-switching, script-switching, instruction following in English, grammatical error correction in English' in addition to the task rests on an implicit decomposition of the task in which the English instruction frame is outside the task. For a benchmark whose input is exactly the English template plus a target-language question, processing that template is part of the input–output mapping. The paper should define 'the task' at the level of abstraction it assumes; otherwise the phrase 'more than just the task' is a terminological claim about where the task boundary is drawn, not an argument. This matters because the recommendation to move away from mixed prompts depends on that boundary.","section":"§3, AfriMMLU example"},{"comment":"The sentence 'Without target language words or (to an extent) language-agnostic labels, the evaluation method and goal will be inherently imprecise' is too strong without qualification. In a cross-lingual evaluation that fixes the same English label set for all languages, the label representation is constant across conditions and is not a source of cross-lingual imprecision; it may be a confound for measuring MLU, but it is a deliberate part of the interface for a task-performance evaluation. The claim should be restricted to the MLU setting or to comparisons where label language is itself a variable.","section":"§3, representation paragraph"}],"minor_comments":[{"comment":"Figure 1 is informative but is never referenced in the body text; please add a pointer, for example at the end of §2 where the diagram is first discussed.","section":"Figure 1"},{"comment":"The nonstandard definition of MLU is important enough to appear in the main text; as written, a reader who skips footnotes may misunderstand the scope of the argument.","section":"Footnote 2"},{"comment":"The GitHub URL in footnote 6 points to a specific commit; for archival purposes, consider citing the repository version or a DOI rather than an ephemeral blob URL.","section":"Footnote 6"},{"comment":"The terms 'imprecise,' 'task,' and 'MLU' are used as technical terms without explicit definitions in one place; a short 'terminology' paragraph in §2 would improve readability and would also help resolve the major concerns above.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written position paper with a clear conceptual core. The main issue is the scope of the 'imprecise even for task performance' claim; if the authors reframe it as 'imprecise for MLU evaluation unless the interface is explicitly part of the construct,' I would be inclined to accept. The paper fits a position-paper track well and should be publishable after that revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a clearly argued position paper. The central move is to distinguish English as an interface from English as a natural language in multilingual prompting, and to point out that mixed prompts test more than the task or the target language. The examples are well chosen: the MaLa-500 news prompt, AfriMMLU's English shell with target-language questions, and Hendy's translation prompt where English is neither source nor target. The authors are right that calling this 'multilingual understanding' is imprecise when what is actually being evaluated is instruction following in English plus script switching and so on. That is a genuinely useful framing, and the paper is honest about the task-performance perspective.\n\nThe soft spot is the conclusion. The authors claim mixed prompts are 'imprecise or misleading evaluations, even if the ultimate goal was to evaluate and improve task performance.' That does not follow from their own argument. In Section 3, the imprecision is established relative to multilingual language understanding: you are not testing MLU, and you are testing more than the task. But if the evaluation target is task performance in a system that will actually be deployed with English as the prompt language, then the English shell is not a confound; it is the operational environment. The paper concedes in Section 2 that task performance is widespread and practical, but never refutes it. It just calls it 'seemingly at odds' with the NLU label. So the recommendation to abandon mixed prompts depends on an unargued normative priority of language understanding over task performance. That is a legitimate position to take, but the paper should say it is advocating for that priority rather than presenting the conclusion as technical necessity.\n\nMinor points: the 'grammatical error correction' bullet relies on a typo in AfriMMLU's prompt, which is an artifact of that benchmark, not a structural feature of all mixed prompts. And the translation naturalness examples are persuasive but depend on a particular notion of natural code-switching. These are fixable.\n\nWho is this for? Anyone building or reading multilingual benchmarks. It deserves a serious referee, but the referee should push back on the overclaim. If the authors qualify the conclusion or defend the normative priority, this could be a solid contribution to the methodological literature.","headline":"A useful two-role framing of English in multilingual eval, but the blanket 'inherently imprecise' conclusion overreaches for task-performance goals.","tokens_in":9064,"tokens_out":2168,"would_cite":true,"duration_ms":22615,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using English as an interface in mixed-language prompts evaluates more than the intended task and is therefore imprecise for measuring multilingual understanding, the paper argues.","keywords":["multilingual language models","evaluation methodology","prompting","mixed prompts","language understanding","task performance","code-switching","English as interface"],"falsifier":"A controlled benchmark comparing mixed-prompts with fully native target-language prompts, matched for formatting and label representation, that finds no systematic score gap would undercut the claim that mixed-prompts add extraneous factors, because the interface could then be considered as clean as native prompts.","tokens_in":8193,"feed_emoji":"🌐","tokens_out":7977,"duration_ms":70040,"temperature":0.7,"pith_summary":"This position paper argues that English plays two distinct roles in multilingual language-model evaluation: as an interface, which is a prompt language used to make models perform better, and as a natural language, which is part of what is being understood. It claims that mixing English with a target language in \"mixed-prompts\" tests more than the intended task or the model's multilingual understanding, adding factors like instruction-following in English, unnatural language or script switching, and even accidental grammatical errors. The paper therefore recommends moving away from English-as-interface prompting and toward native-language or natural code-switched prompts, with language understanding as the goal.","feed_headline":"Mixed-language prompts mislead multilingual model tests","feed_subtitle":"A position paper argues that English-as-interface scores add accidental extras, not language understanding.","key_machinery":"The central object is the distinction between two roles of English in multilingual evaluation, and the named \"mixed-prompt\" as the setup that conflates them. A mixed-prompt is an English instruction or label frame interleaved with target-language content, unlike natural code-switching. The argument works by showing that this setup loads multiple uncontrolled factors onto a single score: label token representations, instruction-following in English, script switching, and unnatural language switching, all in addition to the target-language task. The paper also uses the programming-language analogy to show that English-as-interface would only be clean if prompt meaning were language-agnostic, which it is not.","core_discovery":"The paper's central claim is that using English as an interface in mixed-prompts evaluates more than just the task or multilingual understanding, and is therefore imprecise or misleading. The authors distinguish English as interface, whose goal is task performance, from English as natural language, whose goal is language understanding. A mixed-prompt such as MaLa-500's \"The topic of the news {sentence} is {topic}\" embeds a target-language sentence inside an English frame; with English words as labels, the model is simultaneously evaluated on English instruction following, script switching for non-Latin scripts, the unnaturalness of the switch, and the task itself. Because English is a natural language, not a programming language, these extra factors cannot be separated from task performance. The conclusion follows: \"This all results in imprecise or misleading evaluations, even if the ultimate goal was to evaluate and improve task performance.\"","pith_inferences":["A testable extension the paper does not run is to quantify the confound: compare mixed-prompt and fully native-prompt scores on the same tasks and regress the gap on an English instruction-following benchmark; the paper's position predicts a systematic, not incidental, gap.","The argument generalizes beyond English: any high-resource language used as a bridge for low-resource languages would carry the same interface/natural-language ambiguity, so the critique applies to evaluation schemes that translate into a pivot language.","If mixed-prompts are imprecise at evaluation time, the same imprecision likely transfers to training regimes that fine-tune on mixed-prompt data, implying the recommendation to move away from English-as-interface has consequences for data creation, not just for evaluation."],"forward_implications":["Model rankings from benchmarks that embed target-language sentences in English frames and use English label words should not be read as rankings of multilingual understanding.","Translation evaluations that prompt in English when English is neither source nor target measure prompt-following and interface compatibility in addition to translation quality.","Adopting native-language or natural code-switched prompts will require more instruction-tuning data in target languages, shifting the bottleneck from prompt design to data and modeling.","Reporting \"multilingual understanding\" from mixed-prompt results should be replaced with a description that the model understands English instructions interleaved with target-language passages."],"supporting_citations":[{"why":"Supplies the taxonomy of prompt-based multilingual evaluations and the reported finding that models perform better when the task is in English, the premise the paper critiques.","marker":"Zhang et al. (2023)"},{"why":"MaLa-500 serves as the leading example of an English-frame mixed-prompt with English label words on a 205-language topic-classification task.","marker":"Lin et al. (2024)"},{"why":"SIB-200 is the dataset behind the MaLa-500 example, labeled as NLU, which the paper uses to show the mismatch between task framing and evaluation goal.","marker":"Adelani et al. (2024a)"},{"why":"IrokoBench's AfriMMLU subtask provides a mixed-prompt with English instruction and target-language questions, used to enumerate the extra factors being evaluated.","marker":"Adelani et al. (2024b)"},{"why":"The machine-translation prompt is the key example of English as an interface where English is neither source nor target, showing the interface role most starkly.","marker":"Hendy et al. (2023)"},{"why":"Provides the definition of code-switching used to argue that mixed-prompts are unnatural and therefore not a natural-language phenomenon.","marker":"Milroy and Muysken (1995)"},{"why":"Its multilingual-exemplars setup exemplifies English as a natural language rather than an interface, the contrast that frames the paper's recommendation.","marker":"Shi et al. (2022)"}],"fun_headline_variants":["English-as-interface prompts skew multilingual benchmarks","Why mixed English prompts fail multilingual tests","English framing muddles language model evaluations","English-as-interface contaminates multilingual scores","Drop English crutch in multilingual LM evals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that language understanding, not task performance, should be the primary goal of multilingual LM evaluation; if task performance alone is a legitimate goal, mixed-prompts remain a defensible method.","fun_headline_variants_meta":{"raw":{"variants":["English-as-interface prompts skew multilingual benchmarks","Why mixed English prompts fail multilingual tests","English framing muddles language model evaluations","English-as-interface contaminates multilingual scores","Drop English crutch in multilingual LM evals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1435,"prompt_tokens":821,"completion_tokens":614,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":558}},"tokens_in":437,"tokens_out":614,"duration_ms":6280,"temperature":1.0,"reasoning_tokens":558,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:51:14.120252+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled benchmark comparing mixed-prompts with fully native target-language prompts, matched for formatting and label representation, that finds no systematic score gap would undercut the claim that mixed-prompts add extraneous factors, because the interface could then be considered as clean as native prompts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the taxonomy of prompt-based multilingual evaluations and the reported finding that models perform better when the task is in English, the premise the paper critiques."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the definition of code-switching used to argue that mixed-prompts are unnatural and therefore not a natural-language phenomenon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Its multilingual-exemplars setup exemplifies English as a natural language rather than an interface, the contrast that frames the paper's recommendation."}],"review_version":1}