{"id":"ba6dc7d9-1bdf-439d-84fc-69cd20577b49","arxiv_id":"2412.01186","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SailCompass provides a reproducible benchmark and robustness analysis for evaluating base LLMs on three Southeast Asian languages.","lead":"SailCompass is a new evaluation suite for large language models on Indonesian, Vietnamese, and Thai, covering eight tasks with 14 existing datasets. Its experiments suggest SEA-specialized models still lead general ones, and that balanced pretraining data and careful prompt design matter.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'balanced language distribution' finding in §4.1/§4.3 is not established: the compared models differ in base architecture, tokenizer, training budget, and even chat-tuning status, so attributing results to language-mix is confounded.","rationale":"I read the paper in good faith and acknowledge its concrete contribution: a public benchmark, a reproducible OpenCompass-based harness, and a detailed analysis of MCQ prompt sensitivity. Those contributions do not depend on the questionable causal claim. However, the central finding about balanced language distribution is load-bearing for the abstract and for the paper's practical recommendation, and the published evidence cannot identify language-mix as the cause. The reader's weakest-assumption analysis pinpoints exactly this confound in §4.1 and §4.3, and I agree with it. This is not a disagreement with external consensus; it is an internal identification problem: the claim requires a controlled comparison that the paper does not provide. A single controlled CFT experiment with fixed base model and fixed budget would settle the issue. Because the concern is addressable and the benchmark itself remains valuable, I do not recommend changing the reader's conditional verdict.","tokens_in":19595,"tokens_out":5001,"duration_ms":48844,"concrete_test":"Run a controlled continual-pretraining experiment from a single base model (e.g., Qwen-1.5-7B): condition A continues pretraining on a fixed token budget with a balanced Indonesian/Vietnamese/Thai mix; condition B continues on the same budget with Thai-only or Vietnamese-only data, holding data quality, optimizer, and compute fixed. Evaluate both on the SailCompass FLORES-200 (En↔id/th/vi) and THAI SUM/INDO SUM/XLSUM sets. If condition A does not show a smaller language-degeneration gap and better cross-language summarization than condition B, the §4.1/§4.3 balanced-distribution conclusion fails to land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's second headline finding—that a balanced language distribution is important for developing better SEA-specialized LLMs—is not supported by the comparisons offered. In §4.1, Sailor, SeaLLM, and Sea-Lion are grouped as 'balanced' and observed to have smaller En↔X Chrf++ differences, but these models come from different labs, use different base architectures (Qwen-1.5, Llama-2-family, and from-scratch), different tokenizers, and unreported compute/token budgets, so the smaller gap is equally attributable to any of those factors. In §4.3 the analogous claim that 'monolingual-specific continual pretraining greatly hurts the models' multilingual performance' rests on exactly two models, Typhoon (Llama-3-based) and VinaLLaMA (Llama-2-based), compared against different general baselines, again without controlling for base model or training budget. The protocol is also not clean: footnote 7 admits SeaLLM is evaluated as SeaLLM-Hybrid, an instruction-tuned variant, because the base model is unavailable, placing a post-trained model in the 'base model' comparison. No training-data composition is reported for any 'balanced' model, so the causal attribution to language-mix is unidentifiable from the published results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SailCompass, an evaluation benchmark for Indonesian, Vietnamese, and Thai built on OpenCompass. It covers eight primary tasks and fourteen datasets across generation, multiple-choice, and classification task types, and evaluates open base models under few-shot prompting. The authors investigate prompt variants and perplexity-based ranking for MCQ tasks, use contextual calibration for classification, and report three headline findings: SEA-specialized LLMs still outperform general LLMs with a narrowed gap, a balanced language distribution is important for developing better SEA LLMs, and advanced prompting techniques are necessary for robust evaluation. All datasets and scripts are released publicly.","tokens_in":1622,"tokens_out":2824,"duration_ms":91048,"significance":"If its claims are properly supported, SailCompass would be a useful community resource for SEA-language evaluation: the dataset selection draws on native corpora, the code is built on a widely used framework, and the MCQ prompt-robustness experiments address a real evaluation pitfall. The paper is also commendably transparent in its Limitations section about the base-model-only scope and the three-language coverage. However, one of the two main empirical findings about language balance is confounded by architecture and training differences, and a key quantitative claim about manipulated training is not derivable from the presented table. The benchmark itself is a solid contribution, but the headline findings need either stronger controlled evidence or more cautious framing.","major_comments":[{"comment":"The causal claim that a balanced language distribution is important for SEA-specialized models is not identified by the comparisons in §4.1 and §4.3. The models grouped as 'balanced' (Sailor, SeaLLM, Sea-Lion) differ from Typhoon and VinaLLaMA in base architecture, tokenizer, training data scale, and post-training status, and footnote 7 admits that SeaLLM is represented by the instruction-tuned SeaLLM-Hybrid variant. In §4.3, the conclusion that 'monolingual-specific continual pretraining greatly hurts the models' multilingual performance' rests on exactly two models (Typhoon and VinaLLaMA) against different general baselines, with no training-data composition or compute budget reported for any of the 'balanced' models. Under these conditions the smaller En↔X Chrf++ gaps and summarization scores cannot be uniquely attributed to language mix. I ask the authors to either add controlled continual-pretraining comparisons on a fixed base model with varied language mixes, or relabel the finding as an observational correlation and remove the causal wording from the abstract and conclusion.","section":"§4.1, §4.3, Table 2, footnote 7"},{"comment":"The claim that manipulated training yields a 'significant 17.2% improvement' with prompt LiTiLo is not verifiable from Table 4. The table reports rows for QWEN-1.5-7B and QWEN-1.5-7B M, while Appendix C says the manipulated-training experiment was run on Sailor-7B; no aggregation rule is given that produces 17.2%, and the accompanying statement that 'To performance decreases by about 1.8%' also does not follow from the displayed Exact Match values. The authors should state which model was fine-tuned, report the per-dataset and aggregate numbers used for the percentage change, and reconcile the table labels with Appendix C.","section":"§5.2, Table 4, Appendix C"},{"comment":"The finding that 'English Prompt is Better Than Native Prompt' is not supported by any displayed data. Figure 1 plots Chrf++ and BLEU by translation direction only and does not contrast English-prompt with native-prompt conditions, yet the text uses this finding to justify reporting native-prompt results in later sections. The underlying comparison should be shown, or the claim and the justification should be removed.","section":"§4.1, Figure 1"},{"comment":"The effect of contextual calibration on the classification results is not quantified. Table 5 is labeled only as Exact Match, without stating whether the numbers are calibrated or uncalibrated, and no paired before/after accuracy comparison is provided; Appendix D gives label-count distributions but not Exact Match or F1 scores. The conclusion that calibration improves the faithfulness of classification should be backed by a quantitative comparison of calibrated versus uncalibrated task performance, or the conclusion should be restricted to the label-distribution observation.","section":"§6, Table 5, Appendix D"},{"comment":"The evaluation protocol randomly selects a small number of few-shot examples, but no random seeds or variance information are reported anywhere in the paper. Because the benchmark is advertised as reproducible and robust, a reader cannot rerun the exact evaluations or know the sensitivity of the reported scores to the chosen demonstrations; archiving the seeds used and ideally reporting multiple few-shot draws would resolve this.","section":"§3.3 and §3.4"}],"minor_comments":[{"comment":"There are frequent typos and formatting artifacts in the text (e.g., 'Sou theast', 'formul ations', 'categoried', 'T o' in §4.1); a careful proofread is needed.","section":"Throughout"},{"comment":"The task name 'WISESENTI' in Table 5 contradicts 'WISESIGHT' in Table 1 and the surrounding text; use one consistent name.","section":"Table 1 and Table 5"},{"comment":"Table 2 lists 'BLOOM-7B1' while §3.5 and the model list use 'BLOOM'; unify the notation.","section":"Table 2"},{"comment":"The M3Exam Indonesian column is actually the Javanese split; this should be stated in the table caption or main text, not only in a footnote, because Javanese is a distinct language from Indonesian and the current label overstates Indonesian coverage.","section":"Table 1, footnote 6"},{"comment":"The Limitations section already acknowledges the three-language and base-model scope; this is useful transparency, but the introduction and abstract should avoid implying broader coverage than the benchmark actually provides.","section":"Limitations"},{"comment":"Reference [24] is used for the GPT-3.5-Turbo MT results; please verify that the citation points to the exact experimental setup used for those numbers, including prompt language and few-shot count.","section":"§4.1, references"}],"recommendation":"major_revision","confidential_remarks":"The authors are affiliated with Sea AI Lab, and their own Sailor-7B is among the best-performing models in the benchmark. This is not circularity because the benchmark is external to the model and the scores are transparent, but the editor should ensure the affiliation and potential conflict are clearly disclosed in the camera-ready version. The benchmark contribution is substantial; the revision should focus on reframing the balanced-distribution finding, verifying the numerical claims in Section 5.2, and documenting the random seeds and prompt-language comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a usable evaluation harness for Indonesian, Vietnamese, and Thai, with code and data public, and the prompt-robustness analysis is the real contribution. The headline finding about balanced language distribution is not supported by the comparisons offered; treat it as a hypothesis, not a result.\n\nThe paper aggregates 14 existing datasets into 8 tasks across three languages, and builds on OpenCompass so the harness is reproducible. That is genuinely useful for the SEA NLP community. The MCQ prompt-variant study is the freshest part: it tests five prompt formats, compares generative vs PPL-based scoring, and shows that output-format choice can swing scores by 20+ points (e.g., LiTiLoPPL vs ToPPL for Mistral and Sailor on Belebele). The manipulation experiment on Sailor-7B is also a nice stress test of whether a prompt format can be gamed.\n\nSoft spots, in order of severity. First, the 'balanced language distribution' claim in §4.1 and §4.3 is confounded. Sailor, SeaLLM, Sea-Lion come from different labs with different base architectures, tokenizers, and compute budgets. SeaLLM is evaluated as SeaLLM-Hybrid, an instruction-tuned model, not the base model (footnote 7). So the smaller translation gap or better summarization cannot be attributed to language mix. This is a causal claim with no controlled comparison. Second, the M3Exam 'Indonesian' split is actually Javanese (footnote 6). That means the Indonesian exam numbers are not Indonesian; it's a language mismatch that should be fixed or relabeled. Third, there are no multiple seeds or error bars; single greedy runs. For a benchmark claiming to be 'robust,' that's thin, though for LLM eval it's common practice. Finally, the 17.2% improvement in §5.2 is not derivable from Table 4; the Δavg values there are 0.07, 0.14, 0.07 for M3Exam, not 17.2%. That looks like an inconsistency in reporting, possibly a different table.\n\nWho is this for? People building or evaluating SEA LLMs; they get a ready-made harness and a cautionary study on MCQ prompt sensitivity. The balanced-language conclusion should not be cited without caveats. I'd send it to peer review, but with the expectation of major revision: the confound needs to be acknowledged or re-analyzed, the Javanese issue fixed, and the numbers reconciled.","headline":"A genuinely useful and reproducible SEA evaluation harness with a solid prompt-robustness study, but the headline balanced-language-distribution finding is confounded and the Indonesian exam split is mislabeled.","tokens_in":20366,"tokens_out":3069,"would_cite":true,"duration_ms":24185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SailCompass: SEA-specialized LLMs still lead, but the gap is narrowing.","keywords":["Southeast Asian languages","LLM evaluation","multilingual benchmark","prompt robustness","contextual calibration","continual pretraining","machine translation","multiple-choice evaluation"],"falsifier":"Train the same base model on the same corpus and compute budget with only the SEA language mix varied, for example a balanced Indonesian, Vietnamese, and Thai mixture versus a Thai-only mixture, and compare on SailCompass tasks; if the balanced model does not clearly beat the monolingual one, the paper's central finding fails.","tokens_in":19432,"feed_emoji":"🧭","tokens_out":7184,"duration_ms":62533,"temperature":0.7,"pith_summary":"SailCompass is an evaluation benchmark for Indonesian, Vietnamese, and Thai that combines 14 datasets across eight tasks, including translation, summarization, question answering, exams, commonsense reasoning, reading comprehension, natural language inference, and sentiment analysis. The paper uses it to compare general multilingual LLMs with models specialized for Southeast Asian languages, and argues that the specialized models still perform better overall, though the margin has shrunk. A second claim is that a balanced distribution of SEA languages in the training corpus matters more than single-language focus, because models trained on an even mix of languages show less translation degeneration and better summarization across languages. The paper also shows that evaluation setup is decisive: multiple-choice scores shift sharply with prompt format, and calibration is needed to keep classification tasks from collapsing onto a single label.","feed_headline":"SEA-specialized LLMs still beat general models; gap narrows","feed_subtitle":"A 14-dataset benchmark for Indonesian, Vietnamese, and Thai shows balanced language mix matters most.","key_machinery":"The load-bearing object is the evaluation protocol itself. SailCompass selects datasets built on native corpora where possible, writes task instructions in the target language, uses greedy decoding, and evaluates generation with BLEU and ChrF++. For multiple-choice tasks it varies five prompt configurations, which combine option text, option labels, and output format, and scores answers either by generation or by perplexity-based ranking, appending each option to the prompt and choosing the lowest-perplexity continuation. For classification it applies contextual calibration, which estimates label bias on context-free inputs and renormalizes label scores to counter it. These choices are what let the paper attribute score differences to model ability rather than to prompt artifacts.","core_discovery":"The paper's central discovery is that the competitive landscape for Southeast Asian languages has not flipped: SEA-specialized base LLMs still outscore general multilingual LLMs on SailCompass, but the advantage is smaller than earlier benchmarks suggested. The reason, the paper argues, is visible in translation and summarization: models continually pretrained on a balanced mix of Indonesian, Vietnamese, and Thai lose less cross-lingual ability than models trained mostly on one language, and they summarize other SEA languages better. A third finding is methodological: multiple-choice results depend strongly on whether the model is asked to generate the answer text or an option ID and on how options appear in the prompt, so the paper recommends a configuration that returns answer text scored by perplexity; classification tasks require contextual calibration because without it models fixate on one or two labels.","pith_inferences":["The balanced-language-distribution conclusion is observational, since the compared models come from different labs with different base architectures and training budgets; a controlled experiment that varies only the language mix on one base model could confirm the causal story.","The finding that translated QA benchmarks produce lower scores than a native benchmark suggests translated test sets may underestimate SEA-language ability, a caveat that generalizes to other low-resource languages.","The recommendation to score multiple-choice questions by answer-text perplexity likely transfers to other multilingual evaluation settings, not just Southeast Asian languages.","Because SailCompass evaluates only base models, the findings may not carry over to instruction-tuned chat models, whose prompt sensitivity and label bias differ."],"forward_implications":["Developers building applications for Indonesian, Vietnamese, and Thai should still prefer SEA-specialized models over general LLMs, but should expect the performance gap to keep shrinking.","Continual pretraining for Southeast Asian languages should use a balanced mixture of SEA languages rather than concentrating on one language, because balanced models generalize to the other languages while monolingual models do not.","Multiple-choice evaluations of multilingual LLMs should let the model output the answer text and rank candidates by perplexity, since option-ID prompts bias predictions toward particular labels.","Classification evaluations should apply contextual calibration; without it, models can score near random because they collapse onto one label."],"supporting_citations":[{"why":"Introduces Sailor, a SEA-specialized model whose balanced language mix is a key comparison point.","marker":"[14]"},{"why":"Introduces SeaLLM, another SEA-specialized model used in the continual-pretraining comparisons.","marker":"[27]"},{"why":"Supplies Sea-Lion, the SEA-specialized model trained from scratch, anchoring the cross-training-paradigm comparison.","marker":"[2]"},{"why":"Provides the open evaluation platform that SailCompass builds on for reproducible execution.","marker":"[28]"},{"why":"Supplies the FLORES-200 translation test data and the passages used in the reading comprehension task.","marker":"[13]"},{"why":"Provides the contextual calibration method applied to classification tasks.","marker":"[47]"},{"why":"Documents option-ID bias in multiple-choice selection, motivating the prompt-configuration analysis.","marker":"[48]"},{"why":"Supplies the native Indonesian TyDiQA data used to compare native and translated benchmarks.","marker":"[11]"}],"fun_headline_variants":["SEA-specialized LLMs still beat general, gap narrows","Balanced SEA data matters most for LLM performance","Benchmark: SEA LLMs lead, but prompting tweaks needed","SailCompass: robust eval shows SEA advantage persists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that balanced language distribution causes the gains assumes the SEA-specialized models differ only in language mix, when in fact they come from different labs with different base architectures, tokenizers, and training budgets.","fun_headline_variants_meta":{"raw":{"variants":["SEA-specialized LLMs still beat general, gap narrows","Balanced SEA data matters most for LLM performance","Benchmark: SEA LLMs lead, but prompting tweaks needed","SailCompass: robust eval shows SEA advantage persists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1404,"prompt_tokens":860,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":474}},"tokens_in":476,"tokens_out":544,"duration_ms":5460,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:35:43.156050+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same base model on the same corpus and compute budget with only the SEA language mix varied, for example a balanced Indonesian, Vietnamese, and Thai mixture versus a Thai-only mixture, and compare on SailCompass tasks; if the balanced model does not clearly beat the monolingual one, the paper's central finding fails.","supporting_citations":[{"cited_title":"Sea-lion (southeast asian languages in on e network): A family of large language models for southeast asia","cited_arxiv_id":null,"evidence_quote":"Supplies Sea-Lion, the SEA-specialized model trained from scratch, anchoring the cross-training-paradigm comparison."},{"cited_title":"Opencompass: A universal e valuation platform for foundation models","cited_arxiv_id":null,"evidence_quote":"Provides the open evaluation platform that SailCompass builds on for reproducible execution."},{"cited_title":"Calibrate before use: Improving few-shot performance of language models","cited_arxiv_id":null,"evidence_quote":"Provides the contextual calibration method applied to classification tasks."},{"cited_title":"Llama 2: Open foundation and ﬁne-tuned chat models, 2023","cited_arxiv_id":null,"evidence_quote":"Documents option-ID bias in multiple-choice selection, motivating the prompt-configuration analysis."},{"cited_title":"Clark, Jennimaria Palomaki, Vitaly Nikola ev, Eunsol Choi, Dan Garrette, Michael Collins, and Tom Kwiatkowski","cited_arxiv_id":null,"evidence_quote":"Supplies the native Indonesian TyDiQA data used to compare native and translated benchmarks."}],"review_version":1}