Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Locally sourced SEA benchmarks separate LLMs more sharply than translated English tests, and the paper argues that real-world multilingual ability should be measured with native content.

desk verdict A genuinely useful pair of SEA benchmarks, but the comparative 'better discernment' claim needs better statistics before it can be trusted. read the letter →

arxiv 2502.06298 v1 pith:FZKYUT5Q submitted 2025-02-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords multilingualbenchmarkingSoutheastAsianlanguagesLLMevaluationtranslationbiasopen-endedinstructionfollowingexam-basedsafetyalignmentreal-worldqueries
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SeaExam and SeaBench are two new benchmarks for Indonesian, Thai, and Vietnamese, built from regional school exams and from multi-turn, open-ended instructions written by native speakers rather than translated from English. The paper's central claim is that these locally sourced questions separate LLMs more cleanly than translated versions of MMLU and MT-Bench: they sit closer to real user queries in embedding space, and model scores spread more widely across both models and languages. A sympathetic reader would take this as evidence that translation-based multilingual benchmarks understate real capability gaps, and that open-ended prompts expose cross-language differences that multiple-choice questions hide. The paper also reports that all nine tested models answer SeaBench safety questions markedly worse than other categories, suggesting weak alignment in local contexts. If the claim holds, benchmark design for multilingual evaluation should prioritize native content over translated content.

What carries the argument

The paper's measurement machinery has three parts. 'Wild Queries,' a reference set of 4,658 real user questions in the three languages, is filtered from large public chat logs and used as the target distribution. Alignment is quantified by cluster distance (C-Dist), the Euclidean distance between centroid embeddings of each benchmark and Wild Queries, computed with a multilingual embedding model; smaller C-Dist means closer to real usage. Discriminability is quantified by the standard deviation of model scores across the nine models or across the three languages, with larger spread taken as better separation. For SeaBench, outputs are graded by an LLM judge, GPT-4o, against native-written reference answers with per-category priority aspects, and the judge's choices are checked against three native linguists.

What would settle it

Compare the same nine models on a translated benchmark that has been matched to SeaExam and SeaBench in question difficulty, topic coverage, and length; if the translated set then shows the same rank ordering with equal or larger spread, the claimed discernment advantage would not stand, and the same test could be run against human preference rankings as the gold standard.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that content provenance changes what a benchmark can measure. SeaExam, with 5,451 multiple-choice questions from official Indonesian, Thai, and Vietnamese exams, and SeaBench, with 300 native-crafted two-turn open-ended tasks across ten categories including 'safety' and 'life', are closer in distribution to a filtered set of 4,658 real-world SEA queries than are MMLU-SEA and MT-bench-SEA, the same benchmarks translated by Google Translate. On those translated baselines, nine 7B-9B models show compressed score differences; on SeaExam and SeaBench the reported spread of model scores is roughly nine percent larger, and SeaBench widens across-language gaps by 6.7 percent on average. The authors attribute the extra sensitivity to open-ended format: with no answer choices to lean on, the model must genuinely command the language and the local context. They further find that all models score worst on the safety category, about one point below the top category on the 1-10 scale, which they read as evidence that alignment has not carried over into multilingual SEA scenarios.

Load-bearing premise

The load-bearing premise is that a wider spread of scores across models or languages is a valid sign that a benchmark is better at telling capabilities apart, rather than a sign of difficulty differences, floor or ceiling effects, or evaluation noise.

Editorial extensions

If this is right

  • Translated multilingual benchmarks understate how much 7B-9B models differ on real Indonesian, Thai, and Vietnamese tasks.
  • Open-ended, native-crafted prompts are the more sensitive instrument for revealing which languages a model genuinely commands.
  • Models tuned for SEA languages show more balanced scores across languages on SeaBench, while generic models fall behind.
  • Safety questions expose a systematic weakness: all nine models score lowest there, so alignment built on English or other high-resource data is not reliably transferring.
  • Benchmark builders should treat translation as a weak substitute for local content when the goal is measuring real multilingual use.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's discrimination claim would be to correlate each benchmark's score differences with an independent gold standard, such as human pairwise preference or downstream task success, rather than with within-benchmark variance.
  • The C-Dist result suggests a general design recipe: mining real user queries to calibrate any new multilingual benchmark, not just for SEA languages, is a cheap way to check whether translated items carry the right cultural content.
  • The safety finding implies a concrete alignment recipe: collect red-teaming prompts from local forums and daily-life scenarios, since translated safety tests likely miss culturally specific harms.
  • The paper's Limitations section notes that only one professional linguist per language adjudicated the agreement study, so no inter-rater agreement is reported; that leaves the claimed judge reliability and the SeaBench score differences partly dependent on a single annotator's preferences.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces two new benchmarks for Southeast Asian languages (Indonesian, Thai, Vietnamese): SeaExam, a multi-subject multiple-choice dataset sourced from real regional exams (5,451 questions), and SeaBench, a multi-turn open-ended dataset (300 questions) authored by native linguists across ten categories, including two new categories 'safety' and 'life'. The authors compare these against translated versions of MMLU and MT-bench, evaluating nine 7–9B instruction-tuned models. They claim that (1) SeaExam and SeaBench are more aligned with real-world user queries (measured by cluster distance to a filtered 'Wild Queries' set), and (2) they 'more effectively discern' model performance, supported by larger cross-model standard deviations (9.3% higher for SeaExam, 8.7% for SeaBench) and by larger cross-language standard deviations for SeaBench. They also report a human evaluation showing GPT-4o judge agreement of about 65% with tie votes and 91% without tie votes, and find uniformly low scores on the safety category.

Significance. The benchmarks themselves are valuable resources. The data construction is credible and carefully documented: real exam papers, native-speaker linguists, explicit cleaning steps, and public release with evaluation code. The Wild Queries dataset is a useful addition. If the central claim about 'more effective discernment' were established, the paper would be an important contribution to multilingual evaluation methodology. However, as it stands, the evidence for that claim is thin: it relies on raw standard deviation differences without statistical justification or significance testing, and the human evaluation has no inter-annotator reliability. The paper is strongest as a resource paper and as a demonstration of how locally sourced queries differ from translated ones; it is weaker as a proof that the new benchmarks have superior discriminative validity.

major comments (4)
  1. [Section 3.2.1, Figure 4] The central claim that SeaExam and SeaBench 'more effectively discern' model performance is supported only by higher standard deviations across nine models, with no significance test, confidence interval, or justification of standard deviation as a measure of discriminative validity. A larger raw SD can reflect floor/ceiling effects, item difficulty, or judge noise (especially for SeaBench scores assigned by GPT-4o), not necessarily better ability separation. With n=9, the reported 9.3% and 8.7% differences could be driven by one or two outlier models. The authors should provide a proper statistical comparison—for example, bootstrap confidence intervals for the SD difference, a permutation test, or an alternative metric such as item response theory discrimination parameters, rank correlation stability, or classification consistency.
  2. [Section 3.2.1, Indonesian results] The paper's own discussion of the Indonesian subset undermines the use of standard deviation as an unproblematic measure of discernment. The text states that SeaExam shows no advantage on Indonesian because model performance is uniformly poor there, resulting in a compressed SD. This is exactly a floor-effect concern: if a harder benchmark compresses scores, the SD decreases even though the benchmark may be more informative. The same mechanism could affect the other language averages. The analysis needs an explicit treatment of item difficulty, score distributions, and floor/ceiling effects before the higher SD on average can be interpreted as better discrimination.
  3. [Section 4 and Limitations] The human evaluation uses a single professional linguist per language, and the Limitations section explicitly concedes that no inter-rater agreement is reported. Since the higher SD for SeaBench (Finding 1) is computed on judge-assigned scores, the extra variance could reflect noise or bias in the GPT-4o judge rather than true signal about model ability. The paper should quantify judge reliability—for example, by having a second annotator on a subset and reporting Cohen's kappa, or by reporting the variance of human scores—and should compare the SD of judge scores against the SD of human scores before attributing the difference to benchmark quality.
  4. [Section 3.1, cluster-distance evidence] The alignment result in Section 3.1—that SeaExam and SeaBench are closer to 'Wild Queries' than translated benchmarks—is largely a design property, because the benchmarks were deliberately constructed to reflect local usage. It is a useful sanity check, but it does not independently validate evaluation quality. The causal step from 'closer distribution to real queries' to 'more effectively discerns LLM performance' is not established. The authors should either provide a direct validation of evaluation quality (e.g., correlation with human judgments or downstream task performance) or soften the claim to 'more representative of local usage' rather than 'effectively discern'.
minor comments (6)
  1. [Section 3.1 title] The section title contains a typo: 'Contructed' should be 'Constructed'.
  2. [Figures 4–6] The y-axis labels read 'Stadndard Deviation', which should be 'Standard Deviation'.
  3. [Appendix B.1] The text 'Nvdia A100 GPUs' should be 'NVIDIA A100 GPUs'.
  4. [Section 4, tie-handling] The explanation of the tie threshold (scores differing by 1 or less) and the baselines 'R = 33.3%' and 'R = 50%' is confusing. It should be clarified, preferably in the main text, why these values are the expected chance agreement and how ties are counted in each setup.
  5. [Section 3.2.3, Figure 6] The conversion of SeaBench scores to 'accuracy' (rate of high-score queries) is defined only in the caption of Figure 6. This definition should be moved to the main text, and the choice of the high-score threshold should be justified.
  6. [Limitations] The Limitations section says the dataset will be 'kept private', which appears to contradict the statement in the footnote that SeaExam and SeaBench are publicly available. This should be clarified to distinguish the released benchmark from future held-out questions.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation; the central claim rests on a construction-validity jump and an unvalidated variance metric, not on a fitted parameter or self-citation chain.

full rationale

The paper's two quantitative supports for 'more effectively discern' are (i) cluster-distance alignment with real Wild Queries and (ii) larger cross-model standard deviations. Neither reduces to the benchmark construction by definition. The alignment result is best read as a manipulation check: because SeaBench was built by native linguists instructed to 'reflect the local users' interests, behavior patterns, cultural content and sensitivities' and SeaExam from local exams, a smaller centroid distance to SEA user queries is the expected consequence of the construction procedure, not an independent proof of evaluation quality. This is a validity limitation, not a circular derivation: the measurement could in principle have failed, and no fitted parameter is reused as a prediction. The discriminative-power argument compares raw standard deviations (9.3% and 8.7%) without significance testing or reliability analysis, and the paper itself notes the Indonesian case where uniformly poor model performance compresses variance; this weakens the inference from variance to 'better discernment' but does not make it circular. Self-citations (M3Exam, SeaLLMs, Liu et al. 2024) appear as construction methodology or as evaluated models, not as load-bearing justifications of the main claim. No uniqueness theorem or ansatz is imported from the authors' prior work. The derivation chain is therefore not equivalent to its inputs by construction; the concerns are about construct validity and statistical support, not circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claims rely on unvalidated methodological assumptions about what makes a benchmark 'better' (standard deviation and centroid distance), plus the representative quality of the Wild Queries and the LLM judge. No new physical or theoretical entities are introduced; the datasets are resources, not postulates.

free parameters (1)
  • Tie threshold in judge agreement = 1 point on a 1-10 scale
    Chosen by the authors to define ties when judge scores differ by 1 or less; affects reported agreement rates, especially the 'without tie votes' columns, but is not derived from data. It is an ad hoc scoring convention.
assumptions (4)
  • domain assumption Standard deviation of scores across models is a valid measure of a benchmark's ability to differentiate model capability.
    Invoked in Section 3.2.1 and Figures 4-6 to conclude SeaExam and SeaBench 'more effectively discern' models. No justification or validation of this metric is given.
  • domain assumption Cluster distance between sentence/entity embedding centroids, using bge-multilingual-gemma2, is a valid proxy for alignment between a benchmark and real-world queries.
    Used in Section 3.1 and Figure 3 to claim SeaExam and SeaBench are more aligned with actual local usage. The Euclidean distance between centroids is an unvalidated summary of distributional similarity.
  • domain assumption Queries in Wild Queries, filtered from LMSYS-Chat-1M and WildChat-1M, are representative of real SEA user queries.
    Used as the ground truth for real-world usage in Section 3.1. These chat logs are biased toward users of specific chatbots and may not represent all SEA populations or languages equally.
  • domain assumption GPT-4o is an accurate judge for open-ended SEA language responses.
    The paper validates GPT-4o against human judges, but with only one linguist per language and no inter-rater agreement analysis. The agreement with ties is about 65%, so the judge is only moderately aligned with human preferences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia." pith.science (2026). https://pith.science/paper/FZKYUT5Q

@misc{pith2026250206298,
  author       = {Pith},
  title        = {Pith review of: SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FZKYUT5Q}},
  note         = {Machine review of arXiv:2502.06298}
}
read the original abstract

This study introduces two novel benchmarks, SeaExam and SeaBench, designed to evaluate the capabilities of Large Language Models (LLMs) in Southeast Asian (SEA) application scenarios. Unlike existing multilingual datasets primarily derived from English translations, these benchmarks are constructed based on real-world scenarios from SEA regions. SeaExam draws from regional educational exams to form a comprehensive dataset that encompasses subjects such as local history and literature. In contrast, SeaBench is crafted around multi-turn, open-ended tasks that reflect daily interactions within SEA communities. Our evaluations demonstrate that SeaExam and SeaBench more effectively discern LLM performance on SEA language tasks compared to their translated benchmarks. This highlights the importance of using real-world queries to assess the multilingual capabilities of LLMs.

Figures

Figures reproduced from arXiv: 2502.06298 by the authors.

Figure 1
Figure 1. Compared with local usage queries in Viet [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Data Examples for the three languages in (a) SeaExam and (b) SeaBench. The correct answer for SeaExam [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Cluster distance between each benchmark and Wild Queries. (a) Cluster distance of entity embeddings [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: (a) Accuracy standard deviation across the [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: (a) Accuracy standard deviation across the [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 8
Figure 8. Figure 8: (a) Entity embedding distribution for Wild Queries, SeaExam, and MMLU-SEA, with each benchmark [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: The ranking correlation for SeaBench between six judges for each language. [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: The prompt for reference-guided single-turn single-answer grading. [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: The prompt for reference-guided multi-turn single-answer grading. [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: The prompt to extract entities from a query . [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Instructions for humans to compare the model performance in (a) turn 1, and (b) turn 2. [PITH_FULL_IMAGE:figures/full_fig_p018_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling Language and Culture for Evaluating Multilingual Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new dual-axis evaluation framework shows multilingual LLMs answer culture-specific questions best when the question language matches the cultural context, with partial neuron-level evidence for the effect.

Reference graph

Works this paper leans on

26 extracted references · 3 canonical work pages · cited by 1 Pith paper

  1. [4]

    ArXiv:2404.03608 [cs]

    Sailor: 9 Open Language Models for South-East Asia. ArXiv:2404.03608 [cs]. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al

  2. [5]

    ArXiv:2407.21783 [cs]

    The Llama 3 Herd of Models. ArXiv:2407.21783 [cs]. Yann Dubois, Balázs Galambosi, Percy Liang, and Tat- sunori B Hashimoto

  3. [6]

    arXiv preprint arXiv:2404.04475

    Length-controlled al- pacaeval: A simple way to debias automatic evalua- tors. arXiv preprint arXiv:2404.04475. Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, et al

  4. [7]

    ArXiv:2406.12793 [cs]

    ChatGLM: A Family of Large Lan- guage Models from GLM-130B to GLM-4 All Tools. ArXiv:2406.12793 [cs]. Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Ka- lika Bali, and Sunayana Sitaram

  5. [9]

    ArXiv:2305.07004 [cs]

    Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross- Lingual-Thought Prompting. ArXiv:2305.07004 [cs]. Kaiyu Huang, Fengran Mo, Hongliang Li, You Li, Yuanchi Zhang, Weijian Yi, Yulong Mao, Jinchen Liu, Yuzhuang Xu, Jinan Xu, Jian-Yun Nie, and Yang Liu

  6. [10]

    ArXiv:2405.10936 [cs]

    A Survey on Large Language Models with Multilingualism: Recent Advances and New Fron- tiers. ArXiv:2405.10936 [cs]. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Re- nard Lavaud, Marie-Anne Lachaux, Pierre Stock...

  7. [11]

    ArXiv:2310.06825 [cs]

    Mistral 7B. ArXiv:2310.06825 [cs]. Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen

  8. [12]

    ArXiv:2304.05613 [cs]

    ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. ArXiv:2304.05613 [cs]. Wei Qi Leong, Jian Gang Ngui, Yosephine Su- santo, Hamsawardhini Rengarajan, Kengatharaiyer Sarveswaran, and William Chandra Tjhi. BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language...

Show all 26 references
  1. [13]

    ArXiv:2406.11939 [cs]

    From Crowdsourced Data to High- Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline. ArXiv:2406.11939 [cs]. Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto

  2. [14]

    ArXiv:2406.04770 [cs]

    WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild. ArXiv:2406.04770 [cs]. Chaoqun Liu, Wenxuan Zhang, Yiran Zhao, Anh Tuan Luu, and Lidong Bing

  3. [15]

    ArXiv:2403.10258 [cs]

    Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models. ArXiv:2403.10258 [cs]. Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, et al

  4. [16]

    ArXiv:2406.10118 [cs]

    SEACrowd: A Multilingual Mul- timodal Data Hub and Benchmark Suite for Southeast Asian Languages. ArXiv:2406.10118 [cs]. Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, and Lidong Bing

  5. [17]

    ArXiv:2312.00738 [cs]

    SeaLLMs – Large Language Models for Southeast Asia. ArXiv:2312.00738 [cs]. OpenAI

  6. [18]

    ArXiv:2303.08774 [cs]

    GPT-4 Technical Report. ArXiv:2303.08774 [cs]. Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S. Yu

  7. [19]

    ArXiv:2404.04925 [cs]

    Multilingual Large Language Model: A Survey of Resources, Taxonomy and Fron- tiers. ArXiv:2404.04925 [cs]. Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush V osoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Di- panjan Das, and Jason Wei

  8. [21]

    ArXiv:2408.00118 [cs]

    Gemma 2: Im- proving Open Language Models at a Practical Size. ArXiv:2408.00118 [cs]. Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, Ai Ti Aw, and Nancy F. Chen

  9. [22]

    ArXiv:2309.04766 [cs]

    SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning. ArXiv:2309.04766 [cs]. 10 An Yang, Baosong Yang, Binyuan Hui, et al

  10. [23]

    ArXiv:2407.10671 [cs]

    Qwen2 Technical Report. ArXiv:2407.10671 [cs]. Jiahao Ying, Yixin Cao, Yushi Bai, Qianru Sun, Bo Wang, Wei Tang, Zhaojun Ding, Yizhe Yang, Xuanjing Huang, and Shuicheng Yan

  11. [24]

    ArXiv:2306.05179 [cs]

    M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models. ArXiv:2306.05179 [cs]. Wenxuan Zhang, Hou Pong Chan, Yiran Zhao, Mahani Aljunied, Jianyu Wang, Chaoqun Liu, Yue Deng, Zhiqiang Hu, Weiwen Xu, Yew Ken Chia, Xin Li, and Lidong Bing

  12. [25]

    ArXiv:2407.19672 [cs]

    SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages. ArXiv:2407.19672 [cs]. Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing. 2024a. Auto arena of llms: Automating llm evaluations with agent peer-battles and...

  13. [26]

    The categorization follows the practice in M3Exam (Zhang et al., 2023)

    id th vi Total language 628 729 57 1414 math 428 221 276 925 natural-science 524 372 612 1508 social-science 0 804 800 1604 Total 1580 2126 1745 5451 Table 4: Distribution of subject categories by language for SeaExam. The categorization follows the practice in M3Exam (Zhang e...

  14. [2018]

    In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium

    XNLI: Evaluating Cross- lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium. Association for Computational Lin- guistics. Yuntian Deng, Wenting Zhao, Jack Hessel, Xi...

  15. [2021]

    ArXiv:2009.03300 [cs]

    Measuring Massive Multitask Language Un- derstanding. ArXiv:2009.03300 [cs]. Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei

  16. [2022]

    ArXiv:2210.03057 [cs]

    Language Mod- els are Multilingual Chain-of-Thought Reasoners. ArXiv:2210.03057 [cs]. AI Singapore

  17. [2023]

    In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore

    MEGA: Multilingual Evaluation of Generative AI. In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore. Association for Computa- tional Linguistics. Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, Dav...

  18. [2024]

    ArXiv:2405.15032 [cs]

    Aya 23: Open Weight Releases to Further Multilingual Progress. ArXiv:2405.15032 [cs]. Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yi- jia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.