Pith. sign in

REVIEW 1 major objections 7 minor 1 cited by

The Multilingual Divide and Its Impact on Global AI Safety

T0 review · 1 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that a large, self-widening 'language gap' in LLM capability and safety performance leaves most of the world's languages less capable and less protected, creating disparities in global AI safety that policy should address.

desk verdict A useful policy-facing synthesis of the multilingual safety gap, but it overstates the causal role of capability and leans heavily on self-cited Aya work. read the letter →

arxiv 2505.21344 v1 pith:2J6AZGJX submitted 2025-05-27 cs.AI cs.CL

classification cs.AIcs.CL
keywords languagegapmultilingualAIsafetylow-resourcelanguagesLLMevaluationculturalbiaspolicymodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that large language models remain measurably less capable and less safe in the vast majority of the world's languages than in a small set of globally dominant ones, and that this 'language gap' is a genuine global-safety problem rather than a mere convenience issue. It assembles evidence that lower-resource languages receive more harmful and less faithful model outputs, that multilingual prompts can bypass safety guardrails, and that the gap is self-widening because synthetic data and LLM-as-judge evaluation both favor languages that already have capable models. The paper roots the gap in uneven access to data, compute, and research participation, and argues that international safety initiatives have largely ignored language coverage. Drawing on the authors' experience building massively multilingual models, it reports concrete progress levers — combining human-curated with synthetic data, building evaluation sets alongside models, safety context distillation, and model merging — and translates these into policy recommendations on multilingual dataset creation, transparency about language coverage, and funding for non-English AI research.

What carries the argument

The mechanism carrying the growth argument is the vicious cycle of Section 3: high-resource languages benefit from synthetic data and from LLM-as-judge evaluation, while low-resource languages are starved of both, so the capability and safety divide widens even without any degradation of absolute performance. The countervailing mechanism from the authors' model-building work is safety context distillation — training a model to refuse harmful prompts by imitating a teacher model's safe responses — which the paper reports reduced harmful generations by 78–89% across languages. The 'low-resource double bind,' the simultaneous scarcity of data and compute, is the third load-bearing mechanism explaining why the gap persists once it exists.

What would settle it

Take one model family, measure harmful-output rates in several low-resource languages, then invest heavily in multilingual data and safety training for half those languages while leaving the other half untouched: if harm rates do not fall where capability rises — or fall only for the specific harm categories trained on — the paper's central policy premise fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that a large gap remains in LLM capabilities and safety performance for languages beyond a relatively small handful of globally dominant languages, and that this gap creates disparities in global AI safety. On the evidence side, it synthesizes findings that GPT-4 produces more harmful and less instruction-faithful generations in lower-resource languages, that prompts in such languages can be used to subvert safety guardrails, and that non-Latin scripts incur higher tokenization costs. On the mechanism side, it argues the gap grows through a vicious cycle: synthetic data generation and LLM-as-judge evaluation reward languages that already have capable models, while low-resource languages lack both reliable data and trustworthy measurement. On the remedy side, the paper reports from the authors' Aya initiative that combining human-curated with synthetic data, building evaluation sets alongside models, and applying safety context distillation cut harmful generations by 78–89% across languages, and it argues these lessons justify specific policy interventions.

Load-bearing premise

The load-bearing premise is that making a model more capable in a language — through more data, evaluation, and compute — will also make it safer in that language; the cited evidence shows capability and harm move together but does not prove that fixing capability fixes harm.

Editorial extensions

If this is right

  • Funding open multilingual evaluation sets — translated and locally created — would give both researchers and regulators the tools needed to detect and mitigate harms outside English.
  • Requiring model providers to disclose per-language coverage and safety performance would let governments run language-specific evaluations and make cross-model comparisons possible.
  • Proven cross-lingual safety techniques such as safety context distillation can be deployed at scale, since they cut harmful generations by 78–89% in human-judged evaluations without sacrificing output quality.
  • Without intervention, the vicious cycle will keep widening the gap, leaving speakers of low-resource languages with higher costs and greater exposure to unsafe outputs as AI becomes embedded in services.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the cited evidence is correlational, so a natural test of the paper's thesis is whether capability gains in a low-resource language actually reduce harm or merely shift it into subtler failure modes; the policy case would look different under the second outcome.
  • Editorial inference: the tokenization cost asymmetry implies an economic amplification the paper does not fully develop — users of low-resource languages pay more per unit of meaning — which suggests script-aware tokenization as a concrete, testable technical fix.
  • Editorial inference: the harm categories used in the cited evaluations may themselves be Western-defined, so actual harm in local contexts — what the paper's red-teaming work labels 'local' harms — could be undercounted, making the true multilingual safety gap larger than measured.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 7 minor

Summary. This paper is a position paper argument that a large gap exists between a small set of dominant languages and the rest of the world's languages both in LLM capabilities and in safety performance, and that this gap creates disparities in global AI safety. The authors synthesize prior work on the causes and consequences of the language gap, identify barriers such as data scarcity, compute inequities, and limited transparency, and present six 'lessons' from Cohere's Aya initiative covering data collection, evaluation, cross-institutional collaboration, multilingual training, toxicity mitigation, and access. They close with policy recommendations supporting multilingual dataset creation, transparency, and research and development.

Significance. If the central claim holds, the paper draws timely attention to an under-addressed dimension of AI safety and provides concrete, actionable recommendations for policymakers. Its strengths are the accessible synthesis of a broad literature, the inclusion of independently sourced evidence for the safety gap (e.g., Shen et al., Yong et al., Deng et al.), and the grounded, practical lessons from a large-scale multilingual effort with publicly released resources and datasets. The paper presents no new empirical evidence, and its policy logic relies on a causal relationship between capability gaps and safety harms that is only supported by correlation. It is a useful primer, but the recommendations would be more robust if the causal mechanism were explicitly disentangled and the self-promotional quantitative claims were made precise.

major comments (1)
  1. [Sections 4 and 6] The paper's central policy inference—that closing the language capability gap will improve global AI safety—rests on a causal claim that the cited evidence does not establish. Figure 4 from Shen et al. (2024), together with Yong et al. (2023a) and Deng et al. (2024), demonstrates a correlation between language resource level and harmful or jailbroken outputs, but not that weaker capability is the operative cause. An equally plausible mechanism is that safety alignment, trained predominantly on English, fails to transfer to other languages; under that alternative, increasing capability without commensurate multilingual safety alignment could increase the fluency and therefore the harmfulness of unsafe outputs. The paper's own Section 5.5 reports a 78–89% harm reduction from safety context distillation—a safety-specific intervention, not a capability intervention—yet the paper does not analyze the two mechanisms separately. Since Recommendations 1–3 in Section 6 are all motivated by the claim that capability investment will reduce safety harms, this gap is load-bearing. The authors should either present evidence that capability improvements reduce harm in low-resource languages (e.g., intervention studies that increase capability while holding alignment fixed) or explicitly reframe the recommendations to prioritize multilingual safety alignment and red-teaming alongside, rather than as a consequence of, capability expansion.
minor comments (7)
  1. [Section 2, para. 4] The phrase 'asymptom of historical technological use' should read 'a symptom of historical technological use', and the same typo appears later in the same paragraph.
  2. [Section 5.3, 'One of our core recommendations'] The sentence 'One of our core recommendations is to complement Language-parallel evaluation sets have benefits yet should be used with an understanding of their limitations' is grammatically incomplete; it should be rephrased to state what should complement language-parallel evaluation sets.
  3. [Section 5.4, para. 1] The typo 'cultral' in 'historical and cultral references' should be corrected to 'cultural'.
  4. [Section 5.1] The quantitative claims 'doubled coverage of existing languages covered by AI', 'largest ever collection of multilingual, instruction fine-tuning data', and 'outperform proprietary options for a subset of languages' lack precise baseline definitions and comparison sets; since Section 5 is policy-facing, these claims should be either defined precisely with citations to specific tables/figures in the referenced papers or reworded to avoid unverifiable superlatives.
  5. [Section 2, 'Limited transparency'] The sentence 'Mistral only claims to support a handful of languages' should identify which Mistral model is meant, as different releases have different language coverage.
  6. [General] The spelling 'state-of-art' appears inconsistently; use 'state-of-the-art' throughout.
  7. [References] The use of '∀' as an author placeholder in the Masakhane references (∀ et al., 2020a; 2020b) may confuse readers unfamiliar with the convention; a footnote or alternative citation format would improve clarity.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the central language-gap and safety claims rest on external, independently checkable evidence; heavy self-citation in Section 5 is descriptive of the authors' own publicly released Aya work and is not load-bearing for the core argument.

full rationale

The paper is a narrative position paper, not a derivation-based study; it contains no fitted parameters, equations, or quantitative predictions that could reduce to their inputs by construction. The central claim — that LLM capability and safety performance lag for low-resource languages and that this harms global AI safety — is supported by independent, non-Cohere evidence (Shen et al. 2024, Fig. 4; Yong et al. 2023a; Deng et al. 2024; Ranathunga & de Silva 2022; Joshi et al. 2020; Bapna et al. 2022; OECD 2023; Muennighoff et al. 2023; Longpre et al. 2024), so the existence and salience of the gap do not rest on the authors' own work. Section 5 is explicitly framed as 'lessons we have learned as a lab' and reports results on publicly released artifacts (Aya models, Aya dataset, Global-MMLU, Aya Red-teaming, Aya Vision Bench); because these releases are open and externally checkable, citing them is real evidence rather than circular self-support. The policy recommendations in Section 6 would survive removal of Section 5, since they follow from the externally documented data, evaluation, transparency, and compute disparities in Sections 2–3. The weakest inference — that investment in capability levers such as datasets and compute will reduce safety harms — is a causal-validity gap rather than a circular reduction: the quoted evidence establishes correlation between resource level and harmful outputs, and the paper's own safety-context-distillation result (78–89% harm reduction) is a safety-alignment intervention, not a capability intervention; Section 4 itself names the operative mechanism as a dearth of multilingual safety testing and mitigation. Heavy self-citation by Cohere authors about Cohere initiatives is frequent but descriptive, not load-bearing for the core claim; no uniqueness theorem, ansatz-by-citation, definitional equivalence, or fitted-input-renamed-as-prediction pattern is present. A few peripheral self-claims ('Aya is the largest participatory machine learning research initiative to date'; 'Aya 101 doubled coverage...') are uncited, but they do not support the paper's central argument. Score 2 reflects no significant circularity, with minor self-citation noted.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No mathematical derivations or fitted parameters appear. The paper rests on domain assumptions about the causal link between language coverage and safety harms, plus the adequacy of cited evidence.

assumptions (2)
  • domain assumption The cited empirical studies accurately characterize multilingual safety harms (e.g., Shen et al. 2024 showing higher harmful generation rates in low-resource languages; Yong et al. 2023 on jailbreaks).
    The paper leans on these external studies for its core claim that the language gap creates safety disparities. It does not independently reproduce these measurements.
  • ad hoc to paper Current international safety initiatives (Seoul commitments, EU AI Act, etc.) indeed neglect multilingual safety in a way that undermines their effectiveness.
    Section 4 asserts this 'huge oversight' without a systematic audit of those initiatives' documents.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Multilingual Divide and Its Impact on Global AI Safety." pith.science (2026). https://pith.science/paper/2J6AZGJX

@misc{pith2026250521344,
  author       = {Pith},
  title        = {Pith review of: The Multilingual Divide and Its Impact on Global AI Safety},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2J6AZGJX}},
  note         = {Machine review of arXiv:2505.21344}
}
read the original abstract

Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages. This paper provides researchers, policymakers and governance experts with an overview of key challenges to bridging the "language gap" in AI and minimizing safety risks across languages. We provide an analysis of why the language gap in AI exists and grows, and how it creates disparities in global AI safety. We identify barriers to address these challenges, and recommend how those working in policy and governance can help address safety concerns associated with the language gap by supporting multilingual dataset creation, transparency, and research.

Figures

Figures reproduced from arXiv: 2505.21344 by the authors.

Figure 1
Figure 1. Bridging the Multilingual Divide: We scrutinize the reasons for the language gap in AI, and review and recommend concrete steps to bridging it. We highlight that the language gap must involve safety mitigation across languages, and that open challenges remain. often widely overlooked or completely absent in efforts to advance AI safety, which primarily focus on English or monolingual settings, leading to potential s… view at source ↗
Figure 2
Figure 2. The language gap is clearly visible in the availability of textual datasets across two popular sources: HuggingFace and Wikipedia. Circles represent the number of HuggingFace datasets includ￾ing text per size tag and mentioning a given language. Color indicates the number of Wikipedia pages in the same language, for the six most frequent languages and a diverse selection of lower-resource languages (source: Ranathun… view at source ↗
Figure 3
Figure 3. ChatGPT requires a greater number of tokens to encode the same contents across language scripts that are less well resourced (FLORES datasets (Goyal et al., 2021), data from Ahia et al. (2023)). The number in brackets indicates the count of languages encoded in each script. web include other languages (Blevins & Zettlemoyer, 2022; Briakou et al., 2023), so by default, most models include training data for many langu… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Results from Shen et al. (2024): Lower-resource languages have a higher rate of harmful and irrelevant generations by GPT-4 than higher-resource languages. (Zou et al., 2023). viewpoint. This lack of linguistic diversity means that the abstract “concept space” that und…
Figure 5
Figure 5. Figure 5: Of examples in MMLU requiring cultural or regionally-specific knowledge to answer cor￾rectly, the majority are geographically tied to North America and dominated by Western culture (from Singh et al. (2025)) One of the core recommendations is that evaluations should al…
Figure 6
Figure 6. Figure 6: Human ratings of harmfulness in model generations, before (Aya) and after safety mit￾igation (Aya Safe) (Üstün et al., 2024). Safety context distillation drastically reduces the ratio of harmful generations for harmful prompts across languages. multilingual systems, of…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It

    cs.CL 2025-05 accept novelty 6.0 of 10

    LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.

Reference graph

Works this paper leans on

20 extracted references · 4 canonical work pages · cited by 1 Pith paper

  1. [3]

    Not Enough Data? Deep Learning to the Rescue!

    Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.500. URL https://aclanthology.org/2022.acl-long.500/. Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. Not enough data? deep learning to the rescue!, 2019. URL https://arxiv.org/abs/1911.03118. Zachary A...

  2. [6]

    URL https://arxiv.org/abs/2404.07900. Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong, Moham- mad Rifqi Farhansyah, Thant Thiri Maung, Frederikus Hudi, David Anugraha, Muhammad Ravi Shulthan Habibi, Muhammad Reza Qorib, Amit Agarwal, Joseph Marvin Imperial, Hitesh Laxmichand Patel, Vicky Feliren, Bahrul Ilmi Nasution, Manuel Anton...

  3. [7]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al

    URL https://openreview.net/forum?id=vESNKdEMGp. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. ∀, Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbo- ...

  4. [8]

    Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, San- jana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan

    URL http://arxiv.org/abs/2305.10510. Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, San- jana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. The flores-101 eval- uation benchmark for low-resource and multilingual machine translation. 2021. 25 Annika Grützner-Zahn, Federico Gaspari, Maria Giagkou, St...

  5. [9]

    URL https://aclanthology.org/2021.acl-demo

    doi: 10.18653/v1/2021.acl-demo.35. URL https://aclanthology.org/2021.acl-demo. 35. Dirk Hovy and Diyi Yang. The importance of modeling social factors of language: Theory and prac- tice. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (eds.),Proceedi...

  6. [11]

    doi: 10.18653/v1/2024.acl-long.843

    Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.843. URL https://aclanthology.org/2024.acl-long.843/. Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk, and Scott A. Hale. Casteist but not racist? quantifying disparities in large language model bias between india and the west.CoRR, abs/2309.08573, 2023. URLhttps...

  7. [12]

    URL https://aclanthology.org/2021.eacl-mai n.303

    doi: 10.18653/v1/2021.eacl-main.303. URL https://aclanthology.org/2021.eacl-mai n.303. Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback, 2023. URLhttps://arxiv.org/ abs/...

  8. [13]

    Regina Lenart-Gansiniec, Wojciech Czakon, Łukasz Sułkowski, and Jasna Pocek

    URL http://arxiv.org/abs/2306.13840. Regina Lenart-Gansiniec, Wojciech Czakon, Łukasz Sułkowski, and Jasna Pocek. Understanding crowdsourcing in science, 2023. ISSN 1863-6691. URLhttps://doi.org/10.1007/s11846-022 -00602-z. Zihao Li, Yucheng Shi, Zirui Liu, Fan Yang, Ali Payani, Ninghao Liu, and Mengnan Du. Quantifying multilingual performance of large la...

Show all 20 references
  1. [14]

    doi: 10.18653/v1/2022.aacl-main.62

    Association for Computational Linguistics. doi: 10.18653/v1/2022.aacl-main.62. URL https://aclanthology.org/2022.aacl-main.62/. Reuters. Explainer: What is happening between Armenia and Azerbaijan over Nagorno-Karabakh?,

  2. [15]

    Accessed on Jan

    URL https://www.reuters.com/world/what-is-happening-between-armenia-azerb aijan-over-nagorno-karabakh-2023-09-19/ . Accessed on Jan. 17, 2024. Angelika Romanou, Negar Foroutan, Anna Sotnikova, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Zeming Ch...

  3. [17]

    URLhttps://aclanthology.org/2021.tacl-1.51/

    doi: 10.1162/tacl_a_00401. URLhttps://aclanthology.org/2021.tacl-1.51/. Reva Schwartz, Apostol Vassilev, Kristen K. Greene, Lori Perine, Andrew Burt, and Patrick Hall. Towards a standard for identifying and managing bias in artificial intelligence, 2022-03-15 04:03:00

  4. [18]

    Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi

    URL https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=934464. Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. The language barrier: Dissecting safety challenges of LLMs in multilingual contexts...

  5. [19]

    doi: 10.18653/v1/2024.acl-long.845

    Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.845. URL https://aclanthology.org/2024.acl-long.845/. Eva Vanmassenhove, Dimitar Shterionov, and Matthew Gwilliam. Machine translationese: Ef- fects of algorithmic bias on linguistic complexity in machin...

  6. [20]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E

    URL https://arxiv.org/abs/2410.16153. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. InThir...

  7. [200]

    doi: 10.18653/v1/P18-2032

    Association for Computational Linguistics, 2018. doi: 10.18653/v1/P18-2032. URLhttps: //aclanthology.org/P18-2032. Meng Ji, Meng Ji, Pierrette Bouillon, and Mark Seligman.Cultural and Linguistic Bias of Neural Machine Translation Technology, pp. 100–128. Studies in Natural Lan...

  8. [2021]

    URL https://aclanthology.org/2021.find ings-emnlp.282

    doi: 10.18653/v1/2021.findings-emnlp.282. URL https://aclanthology.org/2021.find ings-emnlp.282. Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. Do all languages cost the same? tokenization in the era of commercial lang...

  9. [2022]

    doi: 10.18653/v1/2022.sumeval-1.5

    Association for Computational Linguistics. doi: 10.18653/v1/2022.sumeval-1.5. URL https://aclanthology.org/2022.sumeval-1.5/. Ashish Sunil Agrawal, Barah Fazili, and Preethi Jyothi. Translation errors significantly impact low- resource languages in cross-lingual learning, 2024...

  10. [2023]

    doi: 10.18653/v1/2023.acl-long.524

    Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.524. URL https://aclanthology.org/2023.acl-long.524/. Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata, Bryan Wilie, Rah- mad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vin...

  11. [2024]

    URL https://arxiv.org/abs/2406.17761. Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Se- bastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Fr...

  12. [2025]

    David Romero, Chenyang Lyu, Haryo Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesus-German Ortiz-Barajas, Emilio Cueva, Jinheon Baek, Soyeong Jeong, et al

    URL https://openreview.net/forum?id=k3gCieTXeY. David Romero, Chenyang Lyu, Haryo Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesus-German Ortiz-Barajas, Emilio Cueva, Jinheon Baek, Soyeong Jeong, et al. Cvqa: Culturally-diverse multilingual visual question ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.