Pith. sign in

REVIEW 2 major objections 7 minor 16 references

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

T0 review · 2 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that a culturally localized 30-language benchmark changes measured multilingual instruction-following, with translation-only tests understating accuracy by 7–22%.

desk verdict A genuinely useful 30-language benchmark whose central MT-underestimation claim is contradicted by the paper's own appendix tables. read the letter →

arxiv 2507.11882 v1 pith:3R35MEDN submitted 2025-07-16 cs.CL

classification cs.CL
keywords multilingualinstructionfollowingbenchmarklocalizationIFEvalextensionlow-resourcelanguagesmachine-translatedevaluationculturaladaptationlargelanguagemodelscross-lingual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that instruction-following ability in large language models is not captured by English-only or machine-translated benchmarks, and that a culturally localized multilingual benchmark changes what is measured. It introduces Marco-Bench-MIF, an extension of IFEval to 30 languages with 541 instruction-response pairs per language, built through translation plus human and LLM verification. Evaluating over 20 models, the paper reports a 25–35 percentage-point accuracy gap between high- and low-resource languages, 45–60% gains from model scale that still leave script-specific deficits, and 7–22% underestimation when machine-translated data is used instead of localized data. If correct, the benchmark becomes a standard instrument for multilingual instruction-following evaluation, and the three findings set the baseline for what models currently can and cannot do outside English.

What carries the argument

The load-bearing mechanism is the localization pipeline that turns IFEval's English items into comparable multilingual items: automated translation, then three localization operations (lexical substitution, topical transposition, and pragmatic restructuring), then verification by human reviewers and an LLM consensus vote, with post-processing that targets six known translation failure points such as keyword consistency and Latin-character frequency in non-Latin scripts. The benchmark also localizes the evaluation logic itself, adapting punctuation, response-language checks, multi-section coherence, and constrained-output checks to each language. This machinery is what lets the paper attribute score differences to model capability rather than to translation artifacts.

What would settle it

Re-score a model on paired localized and machine-translated items that differ only in a culturally substituted term; if accuracy does not differ, the 7–22% localization effect would not survive, and the paper's Table 10/11 data already allow this check for the five languages with both variants.

Watch

Extended reading notes

Core claim

The paper's central claim is that Marco-Bench-MIF is a valid, culturally localized multilingual instruction-following benchmark, and that evaluating large language models on it reveals systematic patterns hidden by English-only or translated tests. For each of 30 languages, the benchmark provides 541 instruction-response pairs ranging from single-constraint expressive instructions to multi-constraint content hybrids. A hybrid pipeline translates and culturally adapts the items, changing capitalization requirements for non-Latin scripts, substituting region-specific references, and restructuring pragmatics, then verifies them in two human review rounds with LLM consensus. Across 20+ models, the paper reports that instruction-level accuracy exceeds prompt-level accuracy by 10–20%, that 70B+ models outperform 8B models by 45–60%, that high-resource languages score 75–85% while low-resource languages score 50–60%, and that machine-translated data underestimates accuracy by 7–22% relative to localized data, with the gap largest for Yoruba and complex constraint combinations.

Load-bearing premise

The cross-language comparisons assume the 541 test items are equally difficult and cover the same constraint mix in every language, even though full cultural localization was applied only to five languages and only partially to the remaining 24.

Editorial extensions

If this is right

  • Machine-translated multilingual benchmarks can mis-order models by 7–22 accuracy points, so future evaluations should report localized scores before drawing cross-lingual conclusions.
  • Scaling model size from 8B to 70B+ buys 45–60% accuracy, but low-resource languages and non-Latin scripts remain bottlenecks even for the strongest proprietary models.
  • Because instruction-level accuracy runs 10–20 points above prompt-level accuracy, multi-constraint compositional adherence is a distinct capability worth measuring separately.
  • Multilingual-specialized open models can match or beat larger general-purpose models on responding in a specified language, so targeted multilingual training is a complementary path to scale.
  • Marco-Bench-MIF's 541 localized items per language give an off-the-shelf test for comparing any new model's multilingual instruction-following behavior across 30 languages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 7–22% localization effect is probably not uniform across items: a stratified breakdown by constraint type would show whether cultural substitution matters most for content constraints, while format constraints stay translation-insensitive; the paper's aggregate numbers leave this untested.
  • Cross-language comparisons inherit an item-difficulty assumption: back-translating localized items and measuring per-item difficulty in English could test whether the high/low-resource gap is partly an artifact of differential localization coverage, since only five languages received full cultural localization.
  • The Turkic transfer correlation (Turkish strongly predicting Kazakh) alongside the Slavic exception suggests that shared script and morphology, not just language-family membership, drive cross-lingual instruction-following transfer; that hypothesis is an adjacent claim, not the paper's main result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper introduces Marco-Bench-MIF, a multilingual instruction-following benchmark that extends IFEval to 30 languages with varying degrees of cultural and linguistic localization. The construction pipeline combines machine translation, manual correction, and LLM-assisted verification, and the benchmark is released publicly. The authors evaluate more than 20 proprietary and open LLMs and report three main findings: a 25–35% accuracy gap between high- and low-resource languages, 45–60% gains from model scale with persistent script-specific challenges, and a 7–22% underestimation of accuracy when machine-translated (MT) data is used instead of localized data. The paper also provides category-level analyses of constraint types and case studies of failure modes.

Significance. If the benchmark construction is sound and the findings hold, Marco-Bench-MIF would be a useful addition to multilingual instruction-following evaluation, addressing a real gap in existing benchmarks such as IFEval and Multi-IF. The paper's strengths include a detailed description of the localization pipeline, a large-scale evaluation across many languages and model families, public release of the data, and per-category results that go beyond aggregate scores. However, two load-bearing issues undermine the headline claims: the paper's own appendix contradicts the MT-underestimation result for two of the five parallel languages, and the cross-language comparisons are threatened by inconsistent localization coverage across the 30 languages. These issues need to be resolved before the central claims can be accepted.

major comments (2)
  1. [Abstract and Section 4.3.4 / Tables 10–11] The claim that machine-translated data underestimates accuracy by 7–22% compared with localized data is not supported by the paper's own aggregate numbers. For Spanish, the prompt-level average is 63.27% localized versus 64.58% MT; for Malay it is 56.31% versus 57.36%, and the instruction-level Malay comparison also favors MT (67.10% vs 68.09%). Appendix A.2 states that localized data 'consistently outperforms' MT data across all languages, which is contradicted by these rows. The Yoruba example is a selected case, not a general trend, and the range 7–22% is not derived from any reported aggregate statistic. Please either recompute the aggregate comparison, restrict the claim to the languages and conditions where it actually holds, or explain explicitly why Spanish and Malay are exceptions; as written, the abstract overstates the evidence.
  2. [Section 3.2.1 / Sections 4.3.1–4.3.2] The cross-language comparisons assume that the 541 items per language are comparable in difficulty and coverage, but Section 3.2.1 states that for the 24 languages outside the five fully localized ones, only SC+EC instances receive full cultural localization while SC+CC and MC+EC cases are 'strategically sampled.' If the mix of fully localized versus merely translated items varies across languages, the observed high- versus low-resource gaps and script-specific effects (e.g., the yo/ne/kk gaps in Table 3 and the ar/zh comparisons in Figure 2) could reflect differences in localization coverage rather than genuine model capability differences. Please report the per-language counts of fully localized versus sampled items, and ideally verify that the headline gaps persist when the analysis is restricted to the fully localized subset.
minor comments (7)
  1. [Abstract] The phrase 'an carefully-curated extension' should be 'a carefully curated extension.'
  2. [Section 3.3] The phrase 'multilingual isntruction following' contains a typo: 'isntruction' should be 'instruction.'
  3. [Table 2 caption] The caption contains 'performane,' which should be 'performance.'
  4. [Figure 1] The figure label 'Our M-IFEval Dataset' is inconsistent with the benchmark name 'Marco-Bench-MIF' used throughout the paper.
  5. [Figure 2 and Section 4.2] Figure 2 includes models such as DeepSeek-R1 and DeepSeek-V3 that do not appear in Table 2; please clarify the complete list of evaluated models and whether Figure 2 is based on the same evaluation runs.
  6. [Section 4.3.4] The phrase 'the gap is more larger' should be 'the gap is larger.'
  7. [Appendix A.2] The sentence 'the performance of most LLMs on localized data consistently outperforms MT data across all languages' is internally contradictory ('most' versus 'consistently') and is not supported by Tables 10–11; please rephrase to match the data actually reported.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation found: the benchmark is an externally grounded extension of IFEval scored by rule-based verification, and the paper's central risk is an overstated MT-vs-localized claim that its own tables partly contradict, which is a data-consistency defect rather than a circular step.

full rationale

This paper contains no derivation chain that reduces to its own inputs, so no circular step is identified. The benchmark is a curation of the externally published IFEval (Zhou et al., 2023) using a stated localization pipeline (lexical substitution, topical transposition, pragmatic restructuring, Section 3.2.2), and all reported accuracies come from IFEval-style rule-based verification (Section 4.1), not from the evaluated models' self-judgment or from fitted parameters renamed as predictions. The two prior works by overlapping authors (Liu et al., 2023 and Romero et al., 2024, which include authors Longyue Wang and Chenyang Lyu) are cited only to support the background premise that cultural alignment matters; they are not load-bearing for any reported result, so they do not constitute circularity. The most circular-looking claim, the abstract's finding (3) that 'machine-translated data underestimates accuracy by 7-22% versus localized data,' is genuinely empirical rather than definitional: the paper's own Tables 10 and 11 show the localized/MT gap is not forced by construction because it is negative for Spanish (prompt-level 63.27% localized vs 64.58% MT) and Malay (56.31% vs 57.36%; instruction-level 67.10% vs 68.09%). This is a serious internal-consistency problem for the paper — Appendix A.2 asserts that 'the performance of most LLMs on localized data consistently outperforms MT data across all languages and model scales,' which those same tables contradict — but it is an overstatement in results reporting, not a circular derivation, and the paper's own Limitations section concedes residual localization biases. The cross-language comparability threat from Section 3.2.1 ('strategic sampling' of SC+CC and MC+EC for 24 languages) is a construct-validity confound, not circularity. Accordingly the circularity score remains in the 0-2 band.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No fitted numeric parameters or new theoretical entities are introduced; the contributions are a dataset and evaluation findings. The load-bearing assumptions are about benchmark comparability, verifier validity, and representativeness of the MT baselines.

assumptions (5)
  • domain assumption IFEval is a valid and reliable base benchmark for instruction following.
    The entire benchmark extends IFEval's 541 instruction-response pairs and inherits its verifier logic (Section 3).
  • domain assumption Rule-based verifiers remain equally valid across all 30 languages after localization adaptations.
    Section 3.3 lists four adaptation types, but no per-language validation or calibration of verifier difficulty is reported; if verifier strictness varies, cross-language gaps are confounded.
  • domain assumption The 541 items per language are comparable in difficulty and constraint coverage.
    Section 3.2.1 applies full localization to SC+EC instances and strategic sampling to SC+CC and MC+EC cases for 24 languages, so per-language item composition may differ.
  • domain assumption Machine-translated baselines are representative of prior multilingual benchmark practice.
    Section 3.2.1 creates MT variants for only five languages; the 7-22% underestimation claim is generalized from this subset.
  • domain assumption Human and LLM consensus verification removed localization artifacts.
    The paper claims two-round human review and three-LLM consensus, but no inter-annotator agreement or audit is reported; residual biases are acknowledged in the Limitations section.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models." pith.science (2026). https://pith.science/paper/3R35MEDN

@misc{pith2026250711882,
  author       = {Pith},
  title        = {Pith review of: Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3R35MEDN}},
  note         = {Machine review of arXiv:2507.11882}
}
read the original abstract

Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly monolingual and centered on English or simply machine translated to other languages, limiting their applicability in multilingual contexts. In this paper, we present an carefully-curated extension of IFEval to a localized multilingual version named Marco-Bench-MIF, covering 30 languages with varying levels of localization. Our benchmark addresses linguistic constraints (e.g., modifying capitalization requirements for Chinese) and cultural references (e.g., substituting region-specific company names in prompts) via a hybrid pipeline combining translation with verification. Through comprehensive evaluation of 20+ LLMs on our Marco-Bench-MIF, we found that: (1) 25-35% accuracy gap between high/low-resource languages, (2) model scales largely impact performance by 45-60% yet persists script-specific challenges, and (3) machine-translated data underestimates accuracy by7-22% versus localized data. Our analysis identifies challenges in multilingual instruction following, including keyword consistency preservation and compositional constraint adherence across languages. Our Marco-Bench-MIF is available at https://github.com/AIDC-AI/Marco-Bench-MIF.

Figures

Figures reproduced from arXiv: 2507.11882 by the authors.

Figure 1
Figure 1. The number of examples and average prompt length for each language in our [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Results of LLMs on Arabic (Ar) and Chi￾nese (Zh) data divided by categories. struction understanding can be achieved at that scales. Moreover, proprietary LLMs all exhibit strong performance against open LLMs, which demonstrates the superiority of these models and that there is still large room for improvement for open models. 4.3.2 Results per Language We split the evaluation results by the language in our Marco-Be… view at source ↗
Figure 3
Figure 3. Comparison between the performance (aver [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

16 extracted references · 1 canonical work pages

  1. [4]

    Preprint, arXiv:2407.21783

    The llama 3 herd of models. Preprint, arXiv:2407.21783. Yun He, Di Jin, Chaoqi Wang, Chloe Bi, Karishma Mandyam, Hejia Zhang, Chen Zhu, Ning Li, Tengyu Xu, Hongjiang Lv, Shruti Bhosale, Chenguang Zhu, Karthik Abinav Sankararaman, Eryk Helenowski, Melanie Kambadur, Aditya Tayade, Hao Ma, Han Fang, and Sinong Wang

  2. [5]

    Preprint, arXiv:2410.15553

    Multi-if: Benchmark- ing llms on multi-turn and multilingual instructions following. Preprint, arXiv:2410.15553. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Os- trow, Akila Welihinda, Alan Hayes, Alec Radford, et al

  3. [6]

    arXiv preprint arXiv:2410.21276

    Gpt-4o system card. arXiv preprint arXiv:2410.21276. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guil- laume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix,...

  4. [7]

    Preprint, arXiv:2310.06825

    Mistral 7b. Preprint, arXiv:2310.06825. Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang, Jinsong Su, Shuming Shi, and Zhaopeng Tu

  5. [8]

    arXiv preprint arXiv:2307.02971

    On the cultural gap in text-to-image generation. arXiv preprint arXiv:2307.02971. Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, et al

  6. [9]

    arXiv preprint arXiv:2406.09948

    Blend: A benchmark for llms on everyday knowledge in diverse cultures and languages. arXiv preprint arXiv:2406.09948. OpenAI

  7. [10]

    arXiv preprint arXiv:2303.08774

    GPT-4 Technical Report. arXiv preprint arXiv:2303.08774. Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng...

  8. [12]

    arXiv preprint arXiv:2406.05967

    Cvqa: Culturally-diverse multilingual visual ques- tion answering benchmark. arXiv preprint arXiv:2406.05967. Gemini Team

Show all 16 references
  1. [13]

    Preprint, arXiv:2312.11805

    Gemini: A family of highly capa- ble multimodal models. Preprint, arXiv:2312.11805. Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupati- raju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu,...

  2. [14]

    Preprint, arXiv:2408.00118

    Gemma 2: Improving open language models at a practical size. Preprint, arXiv:2408.00118. Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen

  3. [15]

    arXiv preprint arXiv:2311.07911

    Instruction-following evalu- ation for large language models. arXiv preprint arXiv:2311.07911. Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei- Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Sh...

  4. [16]

    Preprint, arXiv:2402.07827

    Aya model: An instruction finetuned open-access multilingual language model. Preprint, arXiv:2402.07827. A Appendix A.1 Results per Category Format Style: Models exhibit formatting fragility, with comma/quote adherence showing 40-60 point gaps between small and large models. S...

  5. [2020]

    In Ad- vances in Neural Information Processing Systems , volume 33, pages 1877–1901

    Language models are few-shot learners. In Ad- vances in Neural Information Processing Systems , volume 33, pages 1877–1901. Curran Associates, Inc. Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia...

  6. [2023]

    arXiv preprint arXiv:2308.12966

    Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal...

  7. [2024]

    arXiv preprint arXiv:2410.02677

    Culturalbench: a robust, diverse and challenging benchmark on measuring the (lack of) cultural knowledge of llms. arXiv preprint arXiv:2410.02677. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy ...

  8. [2025]

    Preprint, arXiv:2412.15115

    Qwen2.5 technical report. Preprint, arXiv:2412.15115. David Romero, Chenyang Lyu, Haryo Akbarianto Wi- bowo, Teresa Lynn, Injy Hamed, Aditya Nanda Kishore, Aishik Mandal, Alina Dragonetti, Artem Abzaliev, Atnafu Lambebo Tonja, et al

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.