{"id":"72b61ab5-0751-4467-a65c-c39c480d938e","arxiv_id":"2501.08335","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A continued-pretrained and weight-merged Llama-3.1 model improves scores on Chinese and Indonesian benchmarks but shows mixed results on Malay, logic, and other tests.","lead":"Researchers from Singapore's A*STAR released an open-source language model, LLaMA-3-MERaLiON-8B-Instruct, trained further on English, Indonesian, and Chinese and then merged with the instruction-tuned Llama-3.1 model. The model beats Meta's Llama-3.1 on several multilingual benchmarks, but the gains are not consistent across all tests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'exceeding official Llama-3' is contradicted by the paper's own primary baseline: Malay Cross-MMLU and average Cross-LogiQA fall below Meta-Llama-3.1-8B-Instruct, and Singlish is never evaluated.","rationale":"The reader's weakest assumption identifies Malay and Singlish as unsupported, which is correct. My concern strengthens this: the paper's own tables show a direct contradiction with the abstract's universal 'exceeding' claim under the paper's primary baseline, rather than merely a missing transfer argument. The released model and the reported Indonesian and Chinese gains are still credible enough for conditional acceptance, but the condition must include fixing the baseline ambiguity and adding Malay/Singlish evaluation. This is an internal inconsistency, not a dispute with consensus, so no verdict change beyond the reader's CONDITIONAL is needed.","tokens_in":6424,"tokens_out":7314,"duration_ms":58765,"concrete_test":"Reproduce Tables 2 and 3 with the released checkpoint against both Meta-Llama-3.1-8B-Instruct and Meta-Llama-3-8B-Instruct using the same SEA-Eval prompts and decoding settings. Then check whether the abstract's phrase 'official Llama-3 models' is meant to cover the newer 3.1 baseline or only the older Llama-3 baseline; if the 3.1 baseline is intended, the Malay Cross-MMLU and average Cross-LogiQA entries directly falsify the claim. Separately, evaluate the model on a Singlish benchmark (e.g., a Singlish subset of SEA-Eval or a locally curated Singlish QA set) and on Malay Cross-LogiQA, since neither appears in the current evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is load-bearing in a way that is undercut by the paper's own tables, provided the baseline is the one used in the paper's narrative. Section 3 motivates the method by comparing with Meta-Llama-3.1-8B-Instruct, and Tables 2–5 list that baseline. Under that baseline, MERaLiON is worse on Malay Cross-MMLU (0.613 vs 0.647) and on average Cross-LogiQA (0.526 vs 0.537), with a Chinese deficit (0.528 vs 0.585). Thus the abstract's 'exceeding the capabilities of the official Llama-3 models' is contradicted, not merely unsupported. If the intended baseline is instead Meta-Llama-3-8B-Instruct, the claim mostly holds, but the paper never says so, and the title/abstract promise for Malay and Singlish remains untested: no Singlish evaluation appears anywhere, and Malay is evaluated only on Cross-MMLU, with no Malay or Singlish training tokens listed in Section 2. The narrower claim that continued pretraining improves Indonesian and Chinese knowledge may survive, but the four-language cross-lingual claim as stated does not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes MERaLiON-TextLLM, a continued-pretraining and weight-merging recipe applied to Llama-3.1-8B-Base. The released LLaMA-3-MERaLiON-8B-Instruct model is trained on roughly 125B tokens distributed across English (38B), Indonesian (45B), and Chinese (42B), followed by instruction tuning on a ~3M-pair multilingual corpus and merging with Llama-3.1-8B-Instruct weights to preserve instruction following. The paper evaluates the model on Cross-MMLU, Cross-LogiQA, IndoMMLU, and CN-Eval, reporting per-language and aggregate scores against Meta-Llama-3.1-8B-Instruct and several other 7B-9B models, and releases the checkpoint on Hugging Face. The central claim, as stated in the abstract, is that the approach achieves performance improvements across benchmarks in Chinese, Indonesian, Malay, and Singlish, exceeding the official Llama-3 models; the body, however, supports a narrower claim for Indonesian and Chinese on some benchmarks and does not evaluate Singlish at all.","tokens_in":6675,"tokens_out":7257,"duration_ms":56478,"significance":"If the narrow claims survive scrutiny, the work is a useful empirical data point: continued pretraining on ~125B tokens followed by weight merging yields an open 8B model that improves over Llama-3.1-8B-Instruct on IndoMMLU (0.576 vs 0.548) and CN-Eval (0.514 vs 0.457), and on several Cross-MMLU and Cross-LogiQA rows for English, Chinese, and Indonesian. The released checkpoint and the per-language result tables are concrete assets. However, the headline four-language claim is not established by the reported evidence: Singlish appears nowhere in the evaluation, Malay is tested only on two benchmarks with mixed results, the baseline used in the text is outperformed on several rows, and the evaluation protocol is unspecified. The contribution can be credited only after the claims are re-scoped and the protocol is documented.","major_comments":[{"comment":"The abstract's core claim of \"performance improvements across benchmarks in these languages, exceeding the capabilities of the official Llama-3 models\" is contradicted by the paper's own stated baseline. Against Meta-Llama-3.1-8B-Instruct, which Section 3 and Section 4.1 explicitly identify as the baseline, Table 2 reports a Malay Cross-MMLU drop (0.613 vs 0.647) and Table 3 reports a lower average Cross-LogiQA (0.526 vs 0.537), with Chinese (0.528 vs 0.585) and Malay (0.489 vs 0.523) individually lower. The manuscript must either name Meta-Llama-3-8B-Instruct as the intended official baseline, in which case the claim mostly holds, or qualify the abstract and Section 4.1 to the languages and benchmarks on which the improvement is actually observed.","section":"Abstract and Section 4.1, Tables 2-3"},{"comment":"The title and abstract promise cross-lingual capability in Chinese, Indonesian, Malay, and Singlish, but Section 2 lists training tokens only for English (38B), Indonesian (45B), and Chinese (42B), and Section 4 contains no Singlish evaluation at all. Malay is evaluated only on Cross-MMLU and Cross-LogiQA, and even there the results are mixed relative to the stated baseline. The manuscript needs either to add Malay and Singlish data and evaluation, or to narrow the title, abstract, and conclusion to the languages actually trained and tested.","section":"Section 2 and Section 4"},{"comment":"The text states that the merged model \"consistently outperformed both the baseline Llama-3.1-8B-Instruct and our standalone instruction-tuned variant across English, Chinese, and Indonesian test sets,\" but Table 1 shows a Chinese Cross-LogiQA score of 0.528 against the baseline's 0.585. This internal contradiction must be corrected, and the corresponding claim in Section 4.1 about Cross-LogiQA improvements should be restricted to Indonesian and English.","section":"Section 3, Table 1"},{"comment":"No evaluation protocol is reported: the manuscript does not specify the prompt template, few-shot count, decoding parameters, or answer-extraction method used for Cross-MMLU, Cross-LogiQA, IndoMMLU, or CN-Eval, and no variance, confidence interval, or significance test is provided. Several reported advantages are small (for example, English Cross-LogiQA is 0.591 vs 0.585, and Indonesian Cross-LogiQA is 0.494 vs 0.455), so the quantitative comparisons are not reproducible and may not be robust. A protocol subsection and uncertainty estimates are needed before the comparative claims can be accepted.","section":"Section 4, Tables 1-5"},{"comment":"The two headline benchmarks, Cross-MMLU and Cross-LogiQA, are taken from Wang et al. (2024), which is co-authored by four of the present authors (Xin Huang, Bin Wang, Zhengyuan Liu, and Ai Ti Aw). This does not by itself invalidate the measurements, but because the main multilingual claim rests on these two benchmarks, the overlap should be disclosed explicitly and at least one external multilingual benchmark should be added to guard against inadvertent alignment with the authors' own evaluation suite.","section":"Section 4, benchmark descriptions"}],"minor_comments":[{"comment":"The abstract says the model is built on Llama-3-8B-Base, while Section 3 says Llama-3.1-8B-Base; the model name also alternates between \"MERaLiON-LLaMA-3.1-8B-Instruct\" and \"LLaMA-3-MERaLiON-8B-Instruct\" and should be standardized throughout.","section":"Abstract and Section 3"},{"comment":"The benchmark description says Cross-LogiQA provides parallel question sets in English, Chinese, and Indonesian, but Table 3 includes a Malay column; the description and the table should be reconciled.","section":"Section 4, Cross-LogiQA bullet"},{"comment":"The table captions and column headers are inconsistent (for example, the header \"Model Series Model\" in Tables 2-5), and the captions should identify the source of each number and the exact evaluation setting.","section":"Tables 2-5"},{"comment":"The paper refers to a \"series\" and to future directions, but only one model checkpoint is described; the manuscript should clarify whether this is a technical report announcing a first release and which additional variants are planned.","section":"Title and Section 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads more like a technical report than a finished journal article: key training details, merge hyperparameters, and evaluation prompts are missing, and the claims outrun the evidence on Malay and Singlish. The two headline benchmarks come from the authors' own SEA-Eval paper, which is not disqualifying but should be disclosed. If the authors re-scope the claims to Indonesian and Chinese, add a full evaluation protocol with uncertainty estimates, and include at least one external benchmark, the paper could become suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a useful checkpoint for Indonesian and Chinese, and the paper's headline claim about Malay and Singlish does not survive its own tables. The model, LLaMA-3-MERaLiON-8B-Instruct, is Llama-3.1-8B continued-pretrained on 125B tokens of English, Indonesian, and Chinese, then merged with instruct weights. The authors publish the weights and describe the data mix, which is more than many such reports do.\n\nWhat actually works: the model beats Meta-Llama-3.1-8B-Instruct on IndoMMLU (0.576 vs 0.548) and CN-Eval (0.514 vs 0.457), and improves Cross-MMLU for English, Chinese, and Indonesian. That is a credible narrow result and a genuine resource for the SEA NLP community.\n\nThe problems start with the abstract. It claims the model exceeds \"the capabilities of the official Llama-3 models\" across these languages. Against the paper's own primary baseline, Meta-Llama-3.1-8B-Instruct, the model is worse on Malay Cross-MMLU (0.613 vs 0.647), worse on Chinese and Malay Cross-LogiQA (0.528 vs 0.585 and 0.489 vs 0.523), and worse on average Cross-LogiQA (0.526 vs 0.537). The only way the headline holds is if the intended baseline is the older Llama-3-8B-Instruct, which the paper never says.\n\nThe scope problem is just as serious. The title promises Malay and Singlish, but the training data contains no Malay or Singlish tokens—only English, Indonesian, and Chinese. Singlish is never evaluated anywhere. Malay is tested only on Cross-MMLU and Cross-LogiQA, both from SeaEval, which four of the present authors co-authored. That self-citation is not disqualifying, but the numbers should be read with that in mind.\n\nThere are also no error bars, significance tests, or evaluation protocol details (prompt, few-shot, temperature). The reported gains are plausible but not rigorously established.\n\nFor a reader working on Indonesian or Chinese LLM adaptation, this checkpoint is worth a look. For a reader who cares about Malay or Singlish, the paper has nothing to offer yet. If it goes to a venue, it needs major revisions: soften the abstract, evaluate or drop Malay and Singlish, and add evaluation details. It deserves review because the artifact is real and the Indonesian/Chinese results are likely useful. I would not accept as is.","headline":"Useful Indonesian/Chinese checkpoint; the Malay and Singlish headline claim is undercut by the paper's own tables.","tokens_in":7234,"tokens_out":4322,"would_cite":false,"duration_ms":33334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Continued pretraining on 125B tokens of English, Indonesian, and Chinese, followed by weight merging with the instruction-tuned checkpoint, produces an 8B model that the paper reports as beating Meta-Llama-3.1-8B-Instruct on Cross-MMLU…","keywords":["multilingual large language models","continued pretraining","weight merging","cross-lingual understanding","Indonesian","Chinese","Malay","Singlish"],"falsifier":"Run the released model on held-out native Malay and Singlish benchmarks (for example, a Malay MMLU and a Singlish conversational or cultural-knowledge set) and compare with Meta-Llama-3.1-8B-Instruct; if the model does not exceed the baseline there, the title's cross-lingual claim for those languages fails.","tokens_in":6212,"feed_emoji":"🌏","tokens_out":10773,"duration_ms":79438,"temperature":0.7,"pith_summary":"This paper reports an open-weights 8B model built by continued pretraining of Llama-3.1-8B-Base on roughly 125B tokens (English 38B, Indonesian 45B, Chinese 42B) and then merging the adapted weights with Llama-3.1-8B-Instruct. The central claim is that this recipe beats the official Llama-3.1-8B-Instruct on the model's headline benchmarks: Cross-MMLU average, IndoMMLU, and CN-Eval, while keeping English strong. The intended significance is that a relatively cheap two-stage recipe, continued pretraining plus weight merging, can adapt a strong open base model to regional languages without full retraining or extensive instruction tuning. The paper positions the approach as a template for expanding coverage to other underrepresented languages.","feed_headline":"125B tokens and a weight merge beat Llama-3.1-8B-Instruct","feed_subtitle":"A continued-pretrained base merged with the instruct checkpoint raises Chinese, Indonesian, and general-knowledge scores.","key_machinery":"The machinery is a two-stage pipeline: continued pretraining on a balanced multilingual corpus, then weight merging between the adapted base and the official instruct checkpoint. The corpus is split into English (38B tokens), Indonesian (45B tokens), and Chinese (42B tokens), with domain classification, hyperparameter optimization, and replay techniques used to mitigate catastrophic forgetting. Weight merging combines the parameters of the continued-pretrained base with Llama-3.1-8B-Instruct so that the adapted knowledge and instruction-following abilities coexist in one set of weights. The paper's evidence for the mechanism is the contrast between a standalone instruction-tuned model, which underperforms the baseline, and the merged model, which outperforms it.","core_discovery":"On its own terms, the paper establishes that the merged model LLaMA-3-MERaLiON-8B-Instruct outperforms Meta-Llama-3.1-8B-Instruct on Cross-MMLU (average 0.717 vs 0.690), IndoMMLU (0.576 vs 0.548), and CN-Eval (0.514 vs 0.457). The pattern the authors emphasize is that the gains come from continued pretraining on a balanced multilingual corpus, and that weight merging preserves those gains while restoring instruction-following behavior that a standalone instruction-tuned model lacked. The paper's own tables show the gains are concentrated in English, Chinese, and Indonesian on the general-knowledge benchmark, with a mixed pattern on the logic benchmark.","pith_inferences":["Because the training data described in the paper contains no Malay or Singlish corpus, the title's promise for those languages rests on transfer from English, Indonesian, and Chinese; the right test is a held-out native benchmark in each language.","The same recipe could be applied to other underserved languages by swapping the target corpus and re-merging with the official instruct checkpoint, which would test whether the reported gains are specific to this language trio or a general property of the method.","The improvements in the paper's tables are clearest on knowledge-heavy benchmarks; a reader should not extrapolate them to reasoning-heavy benchmarks without checking the per-benchmark columns, since the Cross-LogiQA average is slightly below the baseline."],"forward_implications":["The released checkpoint gives an 8B open model with stronger Chinese and Indonesian knowledge and general-knowledge performance than Meta-Llama-3.1-8B-Instruct.","Continued pretraining plus weight merging is demonstrated as a resource-efficient alternative to full retraining or large-scale instruction tuning for multilingual adaptation.","Balanced multilingual pretraining on English, Indonesian, and Chinese does not sacrifice English ability on Cross-MMLU, suggesting the recipe can be applied without an English regression.","The reported gains on IndoMMLU and CN-Eval show the approach transfers to domain-specific national knowledge benchmarks, not just translated general-knowledge tests."],"supporting_citations":[{"why":"Provides the Llama-3.1-8B base and instruct checkpoints that the continued pretraining and weight merging start from.","marker":"Dubey et al. [2024]"},{"why":"Contributes the Cross-MMLU and Cross-LogiQA benchmarks used for the main cross-lingual evaluation.","marker":"Wang et al. [2024]"},{"why":"Supplies IndoMMLU, the Indonesian exam benchmark where the model reports its clearest gain over the baseline.","marker":"Koto et al. [2023]"},{"why":"One source of CN-Eval through C-Eval, used to measure the model's Chinese knowledge.","marker":"Huang et al. [2023]"},{"why":"The other source of CN-Eval through CMMLU, used to measure the model's Chinese knowledge.","marker":"Li et al. [2024]"},{"why":"Provides SEA-LION models used as Southeast Asian baselines in the comparison tables.","marker":"Singapore [2024]"},{"why":"Provides Gemma 2 models used as additional baselines in the comparison tables.","marker":"Team et al. [2024]"},{"why":"Provides Qwen2 models used as additional baselines in the comparison tables.","marker":"Yang et al. [2024]"},{"why":"Provides SeaLLMs models used as additional baselines in the comparison tables.","marker":"Zhang et al. [2024]"}],"fun_headline_variants":["Merged Llama-3-8B beats Llama-3.1 on Chinese, Indonesian, and more","125B tokens and a weight merge give Llama-3 an edge over Llama-3.1","Continued pretraining plus merge lifts Llama-3-8B on multilingual benchmarks","Weight-merging beats instruct-tuning alone for multilingual Llama-3"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that continued pretraining on English, Indonesian, and Chinese transfers to Malay and Singlish, since no Malay or Singlish training corpus is described and Singlish is not evaluated anywhere in the paper.","fun_headline_variants_meta":{"raw":{"variants":["Merged Llama-3-8B beats Llama-3.1 on Chinese, Indonesian, and more","125B tokens and a weight merge give Llama-3 an edge over Llama-3.1","Continued pretraining plus merge lifts Llama-3-8B on multilingual benchmarks","Weight-merging beats instruct-tuning alone for multilingual Llama-3"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00116,"raw_usage":{"total_tokens":4754,"prompt_tokens":847,"completion_tokens":3907,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":3810}},"tokens_in":463,"tokens_out":3907,"duration_ms":23018,"temperature":1.0,"reasoning_tokens":3810,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:30:45.918753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the released model on held-out native Malay and Singlish benchmarks (for example, a Malay MMLU and a Singlish conversational or cultural-knowledge set) and compare with Meta-Llama-3.1-8B-Instruct; if the model does not exceed the baseline there, the title's cross-lingual claim for those languages fails.","supporting_citations":[],"review_version":1}