REVIEW 5 major objections 6 minor 55 references
SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read SailCompass: SEA-specialized LLMs still lead, but the gap is narrowing.
desk verdict A genuinely useful and reproducible SEA evaluation harness with a solid prompt-robustness study, but the headline balanced-language-distribution finding is confounded and the Indonesian exam split is mislabeled. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the evaluation protocol itself. SailCompass selects datasets built on native corpora where possible, writes task instructions in the target language, uses greedy decoding, and evaluates generation with BLEU and ChrF++. For multiple-choice tasks it varies five prompt configurations, which combine option text, option labels, and output format, and scores answers either by generation or by perplexity-based ranking, appending each option to the prompt and choosing the lowest-perplexity continuation. For classification it applies contextual calibration, which estimates label bias on context-free inputs and renormalizes label scores to counter it. These choices are what let the paper attribute score differences to model ability rather than to prompt artifacts.
What would settle it
Train the same base model on the same corpus and compute budget with only the SEA language mix varied, for example a balanced Indonesian, Vietnamese, and Thai mixture versus a Thai-only mixture, and compare on SailCompass tasks; if the balanced model does not clearly beat the monolingual one, the paper's central finding fails.
Extended reading notes
Core claim
The paper's central discovery is that the competitive landscape for Southeast Asian languages has not flipped: SEA-specialized base LLMs still outscore general multilingual LLMs on SailCompass, but the advantage is smaller than earlier benchmarks suggested. The reason, the paper argues, is visible in translation and summarization: models continually pretrained on a balanced mix of Indonesian, Vietnamese, and Thai lose less cross-lingual ability than models trained mostly on one language, and they summarize other SEA languages better. A third finding is methodological: multiple-choice results depend strongly on whether the model is asked to generate the answer text or an option ID and on how options appear in the prompt, so the paper recommends a configuration that returns answer text scored by perplexity; classification tasks require contextual calibration because without it models fixate on one or two labels.
Load-bearing premise
The claim that balanced language distribution causes the gains assumes the SEA-specialized models differ only in language mix, when in fact they come from different labs with different base architectures, tokenizers, and training budgets.
Editorial extensions
If this is right
- Developers building applications for Indonesian, Vietnamese, and Thai should still prefer SEA-specialized models over general LLMs, but should expect the performance gap to keep shrinking.
- Continual pretraining for Southeast Asian languages should use a balanced mixture of SEA languages rather than concentrating on one language, because balanced models generalize to the other languages while monolingual models do not.
- Multiple-choice evaluations of multilingual LLMs should let the model output the answer text and rank candidates by perplexity, since option-ID prompts bias predictions toward particular labels.
- Classification evaluations should apply contextual calibration; without it, models can score near random because they collapse onto one label.
Reading between the lines
- The balanced-language-distribution conclusion is observational, since the compared models come from different labs with different base architectures and training budgets; a controlled experiment that varies only the language mix on one base model could confirm the causal story.
- The finding that translated QA benchmarks produce lower scores than a native benchmark suggests translated test sets may underestimate SEA-language ability, a caveat that generalizes to other low-resource languages.
- The recommendation to score multiple-choice questions by answer-text perplexity likely transfers to other multilingual evaluation settings, not just Southeast Asian languages.
- Because SailCompass evaluates only base models, the findings may not carry over to instruction-tuned chat models, whose prompt sensitivity and label bias differ.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SailCompass, an evaluation benchmark for Indonesian, Vietnamese, and Thai built on OpenCompass. It covers eight primary tasks and fourteen datasets across generation, multiple-choice, and classification task types, and evaluates open base models under few-shot prompting. The authors investigate prompt variants and perplexity-based ranking for MCQ tasks, use contextual calibration for classification, and report three headline findings: SEA-specialized LLMs still outperform general LLMs with a narrowed gap, a balanced language distribution is important for developing better SEA LLMs, and advanced prompting techniques are necessary for robust evaluation. All datasets and scripts are released publicly.
Significance. If its claims are properly supported, SailCompass would be a useful community resource for SEA-language evaluation: the dataset selection draws on native corpora, the code is built on a widely used framework, and the MCQ prompt-robustness experiments address a real evaluation pitfall. The paper is also commendably transparent in its Limitations section about the base-model-only scope and the three-language coverage. However, one of the two main empirical findings about language balance is confounded by architecture and training differences, and a key quantitative claim about manipulated training is not derivable from the presented table. The benchmark itself is a solid contribution, but the headline findings need either stronger controlled evidence or more cautious framing.
major comments (5)
- [§4.1, §4.3, Table 2, footnote 7] The causal claim that a balanced language distribution is important for SEA-specialized models is not identified by the comparisons in §4.1 and §4.3. The models grouped as 'balanced' (Sailor, SeaLLM, Sea-Lion) differ from Typhoon and VinaLLaMA in base architecture, tokenizer, training data scale, and post-training status, and footnote 7 admits that SeaLLM is represented by the instruction-tuned SeaLLM-Hybrid variant. In §4.3, the conclusion that 'monolingual-specific continual pretraining greatly hurts the models' multilingual performance' rests on exactly two models (Typhoon and VinaLLaMA) against different general baselines, with no training-data composition or compute budget reported for any of the 'balanced' models. Under these conditions the smaller En↔X Chrf++ gaps and summarization scores cannot be uniquely attributed to language mix. I ask the authors to either add controlled continual-pretraining comparisons on a fixed base model with varied language mixes, or relabel the finding as an observational correlation and remove the causal wording from the abstract and conclusion.
- [§5.2, Table 4, Appendix C] The claim that manipulated training yields a 'significant 17.2% improvement' with prompt LiTiLo is not verifiable from Table 4. The table reports rows for QWEN-1.5-7B and QWEN-1.5-7B M, while Appendix C says the manipulated-training experiment was run on Sailor-7B; no aggregation rule is given that produces 17.2%, and the accompanying statement that 'To performance decreases by about 1.8%' also does not follow from the displayed Exact Match values. The authors should state which model was fine-tuned, report the per-dataset and aggregate numbers used for the percentage change, and reconcile the table labels with Appendix C.
- [§4.1, Figure 1] The finding that 'English Prompt is Better Than Native Prompt' is not supported by any displayed data. Figure 1 plots Chrf++ and BLEU by translation direction only and does not contrast English-prompt with native-prompt conditions, yet the text uses this finding to justify reporting native-prompt results in later sections. The underlying comparison should be shown, or the claim and the justification should be removed.
- [§6, Table 5, Appendix D] The effect of contextual calibration on the classification results is not quantified. Table 5 is labeled only as Exact Match, without stating whether the numbers are calibrated or uncalibrated, and no paired before/after accuracy comparison is provided; Appendix D gives label-count distributions but not Exact Match or F1 scores. The conclusion that calibration improves the faithfulness of classification should be backed by a quantitative comparison of calibrated versus uncalibrated task performance, or the conclusion should be restricted to the label-distribution observation.
- [§3.3 and §3.4] The evaluation protocol randomly selects a small number of few-shot examples, but no random seeds or variance information are reported anywhere in the paper. Because the benchmark is advertised as reproducible and robust, a reader cannot rerun the exact evaluations or know the sensitivity of the reported scores to the chosen demonstrations; archiving the seeds used and ideally reporting multiple few-shot draws would resolve this.
minor comments (6)
- [Throughout] There are frequent typos and formatting artifacts in the text (e.g., 'Sou theast', 'formul ations', 'categoried', 'T o' in §4.1); a careful proofread is needed.
- [Table 1 and Table 5] The task name 'WISESENTI' in Table 5 contradicts 'WISESIGHT' in Table 1 and the surrounding text; use one consistent name.
- [Table 2] Table 2 lists 'BLOOM-7B1' while §3.5 and the model list use 'BLOOM'; unify the notation.
- [Table 1, footnote 6] The M3Exam Indonesian column is actually the Javanese split; this should be stated in the table caption or main text, not only in a footnote, because Javanese is a distinct language from Indonesian and the current label overstates Indonesian coverage.
- [Limitations] The Limitations section already acknowledges the three-language and base-model scope; this is useful transparency, but the introduction and abstract should avoid implying broader coverage than the benchmark actually provides.
- [§4.1, references] Reference [24] is used for the GPT-3.5-Turbo MT results; please verify that the citation points to the exact experimental setup used for those numbers, including prompt language and few-shot count.
Circularity Check
No significant circularity: the evaluation results are external measurements, and the causal 'balanced distribution' claim is a confounded inference, not a derivation from its own inputs.
full rationale
SailCompass reports scores of existing open models on public datasets (FLORES-200, XQuAD, TyDiQA, XCOPA, BELEBELE, XNLI, etc.) with fixed prompts and metrics; no model parameter is fitted to SailCompass outcomes, and no equation defines the benchmark score in terms of the conclusions. The paper's first finding ('SEA-specialized LLMs still outperform general LLMs') is a direct empirical comparison, not a construction. The second finding ('balanced language distribution is important') attributes observed MT and summarization differences to training-data language mix, but the compared models differ in base architecture, tokenizer, training budget, and in SeaLLM's case chat-tuning status (footnote 7), so the attribution is confounded; confoundedness is a validity problem, not circularity, because the grouping (general vs. SEA-specialized vs. from-scratch) is asserted from released model descriptions rather than derived from the benchmark scores. Self-citations, notably to the authors' Sailor model [14], are used only to describe the evaluated model and to contextualize results; the benchmark datasets and evaluation scripts are external to Sailor, so the self-citation is not load-bearing. The MCQ prompt-robustness analyses are similarly empirical: the manipulated-training experiment (Appendix C) tests prompt variants against an external MMLU auxiliary split and does not define performance as equal to the chosen configuration. The limitations section candidly lists coverage and method gaps but does not reveal any step where an output is identical by construction to an input. Overall, no circular derivation chain is present.
Assumptions & free parameters
free parameters (2)
- few-shot example count =
3 (1 for summarization)
- MCQ prompt configuration =
ToPPL
assumptions (5)
- domain assumption BLEU and Chrf++ adequately measure generation quality for SEA languages
- domain assumption Exact Match is an appropriate metric for MCQ and classification tasks
- ad hoc to paper Machine-translated instructions preserve task meaning
- domain assumption Observed cross-model differences are attributable to training data language distribution
- ad hoc to paper M3Exam Javanese split counts as Indonesian coverage
Cite this review
Pith. "Pith review of SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages." pith.science (2026). https://pith.science/paper/AUOGDFJG
@misc{pith2026241201186,
author = {Pith},
title = {Pith review of: SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/AUOGDFJG}},
note = {Machine review of arXiv:2412.01186}
}
read the original abstract
In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks including 14 datasets covering three task types (generation, multiple-choice questions, and classification). To improve the robustness of the evaluation approach, we explore different prompt configurations for multiple-choice questions and leverage calibrations to improve the faithfulness of classification tasks. With SailCompass, we derive the following findings: (1) SEA-specialized LLMs still outperform general LLMs, although the gap has narrowed; (2) A balanced language distribution is important for developing better SEA-specialized LLMs; (3) Advanced prompting techniques (e.g., calibration, perplexity-based ranking) are necessary to better utilize LLMs. All datasets and evaluation scripts are public.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
MEGA: multilingual evaluation of gen erative AI
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Oc hieng, Krithika Ramesh, Prachi Jain, Akshay Uttama Nambi, Tanuja Ganu, Sameer Segal, Moham ed Ahmed, Kalika Bali, and Sunayana Sitaram. MEGA: multilingual evaluation of gen erative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Languag e Processing, EMNLP 2023 , pages 4232–...
-
[2]
AI Singapore. Sea-lion (southeast asian languages in on e network): A family of large language models for southeast asia. https://github.com/aisingapore/sealion, 2023
work page 2023
-
[3]
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel C ahyawijaya, Ade Romadhony, Rah- mad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, and Sebastian Ruder. One country, 700+ language s: NLP challenges for under- represented languages and dialects in Indonesia. In Smaran da Muresan, Preslav Nakov, and Al...
-
[4]
On th e cross-lingual trans- ferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Y ogatama. On th e cross-lingual trans- ferability of monolingual representations. In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics, ACL 2 020, Online, July 5- 10, 2020 , pages 4623–4637. Association for Computational Linguist ics, 2020. URL https://doi.org/10.18653/v1/20...
-
[5]
BUFFET: Benchmarking large language models for cross-lingual few-shot transfer
Akari Asai, Sneha Kudugunta, Xinyan V elocity Y u, Terra B levins, Hila B Gonen, Machel Reid, Y ulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajish irzi. BUFFET: Benchmarking large language models for cross-lingual few-shot transfer . In NAACL, 2024. 10
work page 2024
-
[6]
Jinze Bai, Shuai Bai, Y unfei Chu, Zeyu Cui, Kai Dang, Xiao dong Deng, Y ang Fan, Wenbin Ge, Y u Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin , Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingz hang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Ben...
work page 2023
-
[7]
The belebele benchmark: a parallel reading comprehension data set in 122 language variants
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Ar tetxe, Satya Narayan Shukla, Don- ald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoy er, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension data set in 122 language variants. CoRR, abs/2308.16884, 2023. URL https://doi.org/10.48550/arXiv.2308.16884
-
[8]
Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Gent a Indra Winata, Bryan Wilie, Fa- jri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadho ny, Karissa Vincentio, Jen- nifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus H udi, Muhammad Satrio Wicak- sono, Ivan Halim Parmonangan, Ika Alfina, Ilham Firdausi Put ra, Samsul Rahmadani, Y u- lianti ...
2023
Show all 55 references
-
[9]
MTG: A benchmark suite for multilingual text gene ration
Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang, Jing jing Xu, Jiaze Chen, Hao Zhou, and Lei Li. MTG: A benchmark suite for multilingual text gene ration. In Findings of the Association for Computational Linguistics: NAACL 2022 , pages 2508–2527, 2022. URL https://doi.org/1...
2022 doi
-
[10]
Using knowledge distillation from keyword extraction to improve the informativeness of neural cross-lingual summarizatio n
Nakhun Chumpolsathien. Using knowledge distillation from keyword extraction to improve the informativeness of neural cross-lingual summarizatio n. Master’s thesis, Beijing Institute of Technology, 2020
2020
-
[11]
Clark, Jennimaria Palomaki, Vitaly Nikola ev, Eunsol Choi, Dan Garrette, Michael Collins, and Tom Kwiatkowski
Jonathan H. Clark, Jennimaria Palomaki, Vitaly Nikola ev, Eunsol Choi, Dan Garrette, Michael Collins, and Tom Kwiatkowski. Tydi QA: A benchmark for infor mation-seeking question answering in typologically diverse languages. Trans. Assoc. Comput. Linguistics , 8:454–470,
-
[12]
Bowman, Hol- ger Schwenk, and V eselin Stoyanov
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina W illiams, Samuel R. Bowman, Hol- ger Schwenk, and V eselin Stoyanov. XNLI: evaluating cross- lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in N atural Language Process- ing, Br...
2018 doi
-
[14]
Sailor: Open language models for south-east asia
Longxu Dou, Qian Liu, Guangtao Zeng, Jia Guo, Jiahui Zho u, Wei Lu, and Min Lin. Sailor: Open language models for south-east asia. arXiv preprint arXiv:2404.03608 , 2024. 11
2024 arXiv
-
[15]
Saiful Islam, K azi Mubasshir, Y uan-Fang Li, Y ong-Bin Kang, M
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, K azi Mubasshir, Y uan-Fang Li, Y ong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. XL- sum: Large- scale multilingual abstractive summarization for 44 langu ages. In Findings of the Association for Computational Linguisti...
2021
- [16]
-
[17]
Em otion recogni- tion for vietnamese social media text
V ong Anh Ho, Duong Huynh-Cong Nguyen, Danh Hoang Nguyen , Linh Thi-V an Pham, Duc-Vu Nguyen, Kiet V an Nguyen, and Ngan Luu-Thuy Nguyen. Em otion recogni- tion for vietnamese social media text. In Computational Linguistics - 16th Interna- tional Conference of the Pacific Assoc...
2019
-
[18]
XTREME: A massively multilingual multi-task benc hmark for evaluating cross- lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Ne ubig, Orhan Firat, and Melvin Johnson. XTREME: A massively multilingual multi-task benc hmark for evaluating cross- lingual generalization. In Proceedings of the 37th International Conference on Ma- chine Learning, ICML 20...
2020
-
[19]
Measuring massive multitask language un derstanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, M antas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language un derstanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtua l Event, Austria, May 3-7, 2021 . OpenRev...
2021
-
[20]
Indolem and indobert: A benchmark dataset and pre-trained language model for indon esian NLP
Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Bald win. Indolem and indobert: A benchmark dataset and pre-trained language model for indon esian NLP. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online),...
2020 doi
-
[21]
Indosum: A new bench mark dataset for indone- sian text summarization
Kemal Kurniawan and Samuel Louvan. Indosum: A new bench mark dataset for indone- sian text summarization. In 2018 International Conference on Asian Language Processin g, IALP 2018, Bandung, Indonesia, November 15-17, 2018 , pages 215–220. IEEE, 2018. URL https://doi.org/10.110...
2018
-
[22]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensc h, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengy el, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023
-
[23]
XGLUE: A new b enchmark datasetfor cross-lingual pre-training, understanding and generatio n
Y aobo Liang, Nan Duan, Y eyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Y ang, Daniel Campo...
2020
-
[24]
Chain-of-dictionary prompting elicits translation i n large language models
Hongyuan Lu, Haoyang Huang, Dongdong Zhang, Haoran Y an g, Wai Lam, and Furu Wei. Chain-of-dictionary prompting elicits translation i n large language models. ArXiv, abs/2305.06575, 2023
2023 arXiv
-
[25]
BHASA: A ho listic southeast asian linguistic and cultural evaluation suite for large language models
Wei Qi Leong, Jian Gang Ngui, Y osephine Susanto, Hamsaw ardhini Rengarajan, Kengath- araiyer Sarveswaran, and William-Chandra Tjhi. BHASA: A ho listic southeast asian linguistic and cultural evaluation suite for large language models. CoRR, abs/2309.06085, 2023. URL https://d...
-
[26]
Vinallama: Llama-b ased vietnamese foundation model, 2023
Quan Nguyen, Huy Pham, and Dung Dao. Vinallama: Llama-b ased vietnamese foundation model, 2023
2023
-
[27]
URL https://doi.org/10.18653/v1/2020.emnlp-main.484
2020 doi
-
[28]
Opencompass: A universal e valuation platform for foundation models
OpenCompass Contributors. Opencompass: A universal e valuation platform for foundation models. https://github.com/open-compass/opencompass, 2023
2023
-
[29]
Indonli: A natural language inference dataset for in donesian
Rahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahru rrozi Rahman, and Clara V ania. Indonli: A natural language inference dataset for in donesian. In Proceed- ings of the 2021 Conference on Empirical Methods in Natural L anguage Process- ing, EMNLP 2021, Virtual Event / Pun...
2021 doi
-
[30]
Ty- phoon: Thai large language models, 2023
Kunat Pipatanakul, Phatrasek Jirabovonvisut, Potsaw ee Manakul, Sittipong Sripaisarn- mongkol, Ruangsak Patomwong, Pathomporn Chokchainant, an d Kasima Tharnpipitchai. Ty- phoon: Thai large language models, 2023
2023
-
[31]
Seallms - large language models for southeast asia
Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunie d, Qingyu Tan, Liying Cheng, Guanzheng Chen, Y ue Deng, Sen Y ang, Chaoqun Liu, Hang Zhang, and Lidong Bing. Seallms - large language models for southeast asia. CoRR, abs/2312.00738, 2023. URL https://doi.org/10.48550/arXi...
-
[32]
chrF: character n-gram F-score for automatic MT evaluati on
Maja Popovi ´c. chrF: character n-gram F-score for automatic MT evaluati on. In Ond ˇrej Bojar, Rajan Chatterjee, Christian Federmann, Barry Haddo w, Chris Hokamp, Matthias Huck, V arvara Logacheva, and Pavel Pecina, editors, Proceedings of the T enth W ork- shop on Statistica...
-
[33]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jin g Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics ,...
-
[34]
Leveraging large la nguage models for multiple choice question answering
Joshua Robinson and David Wingate. Leveraging large la nguage models for multiple choice question answering. In The Eleventh International Conference on Learning Rep- resentations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. URL https://openreview.net/pdf?...
2023
-
[35]
Quantifying language models’ sensitivity to spurious features in prompt design or: How i l earned to start worrying about prompt formatting
Melanie Sclar, Y ejin Choi, Y ulia Tsvetkov, and Alane Su hr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i l earned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Represen tations,
-
[36]
XCOP A: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavas, Olga Majewska, Qian chu Liu, Ivan Vulic, and Anna Korhonen. XCOP A: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Langu age Processing, EMNLP 2020, Online, Novem...
2020
-
[37]
URL https://doi.org/10.18653/v1/2020.emnlp-main.185
2020 doi
-
[38]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dad ashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdher y, Adam Roberts, Aditya 13 Barua, Alex Botev,...
2024
-
[39]
Spruit, C
Nllb team, Marta Ruiz Costa-jussà, James Cross, Onur cC elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Dani el Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Alison Y oungblood, Bap i Akula, Loïc Barrault, Gabriel Mejia Gonza...
2022 arXiv
-
[40]
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Proces sing, 2016. URL https://aclanthology.org/D16-1264
2016
-
[41]
Alex Wang, Y ada Pruksachatkun, Nikita Nangia, Amanpre et Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A sti ckier bench- mark for general-purpose language understanding systems. In Advances in Neu- ral Information Processing Systems 32: Annua...
2019
-
[42]
Bow- man
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill , Omer Levy, and Samuel R. Bow- man. GLUE: A multi-task benchmark and analysis platform for natural language understand- ing. In 7th International Conference on Learning Representations , ICLR 2019 , 2019. URL https://open...
2019
-
[43]
Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Y ang D ing, Ai Ti A w, and Nancy F. Chen. Seaeval for multilingual foundation models: From cro ss-lingual alignment to cultural reasoning. In NAACL, 2024
2024
-
[44]
Chi, Nathanael Schärli, and Denny Zhou
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, ICML 2023, 2 3-29 July 2023, Honolulu, Hawai...
2023
-
[45]
Pythainlp/wisesight-sentiment: First release, September 2019
Arthit Suriyawongkul, Ekapol Chuangsuwanich, Pattar awat Chormai, and Charin Pol- panumas. Pythainlp/wisesight-sentiment: First release, September 2019. URL https://doi.org/10.5281/zenodo.3457447
2019 doi
-
[46]
M3exam: A multilingual, multimodal, multilevel benc hmark for examin- ing large language models
Wenxuan Zhang, Mahani Aljunied, Chang Gao, Y ew Ken Chia , and Lidong Bing. M3exam: A multilingual, multimodal, multilevel benc hmark for examin- ing large language models. In Advances in Neural Information Processing Sys- tems 36: Annual Conference on Neural Information Proce...
2023
-
[47]
Calibrate before use: Improving few-shot performance of language models
Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, 2021
2021
-
[48]
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Y asmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shrut i Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucuru ll, David Esiobu, Jude Fernan- des, Jeremy ...
2023
-
[52]
my answer is c
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel , Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. "my answer is c": First-token p robabilities do not match text answers in instruction-tuned language models, 2024. 14
2024
-
[53]
BigScience Workshop, :, Teven Le Scao, Angela Fan, Chri stopher Akiki, Ellie Pavlick, Suzana Ili´c, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccion i, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Al bert Webson, Pawan Sasanka Ammanam...
2023
-
[56]
auxiliary train
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minl ie Huang. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=shr9PXz7T0. 16 A Prompt Variants fo...
2024
-
[2002]
Association for Computational Linguistics. doi: 10. 3115/1073083.1073135. URL https://aclanthology.org/P02-1040
-
[2015]
Association for Computational Linguistics. doi: 10. 18653/v1/W15-3049. URL https://aclanthology.org/W15-3049
-
[2020]
URL https://doi.org/10.1162/tacl_a_00317
-
[2022]
Association for Computational Linguistics. doi: 10. 18653/v1/2022.acl-long.500. URL https://aclanthology.org/2022.acl-long.500
2022
-
[2024]
URL https://openreview.net/forum?id=RIu5lyNXjT
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.