{"id":"a9dc7128-bfdd-4435-8a05-e4e8163a580c","arxiv_id":"2412.13860","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Continual pretraining of Llama 3 8B on synthetic Nepali-English data improves its Nepali generation but causes English forgetting, with limited evidence for latent retention.","lead":"This paper tests whether a large language model can be taught Nepali by continuing to train it on machine-translated text. The adapted model generates more fluent Nepali than the base model, but loses accuracy on English benchmarks, and the evidence for the claimed 'latent retention' is weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4o evaluation is unvalidated; without human correlation the central claim that the model 'learned Nepali' is not established.","rationale":"The reader's verdict is CONDITIONAL with the weakest assumption identified as synthetic-data quality. I agree that data quality matters, but I argue the more load-bearing concern is the validity of the GPT-4o evaluation. The paper's Q1 result is the primary evidence for the central claim, and it is entirely mediated by an unvalidated automatic judge. The parallel corpus actually uses organic Nepali (OSCAR) with synthetic English, so the reader's 'translationese' concern applies mainly to the instruction set, not the 5M pretraining pairs. The evaluation issue is more direct: the paper provides no evidence that GPT-4o's Nepali scoring correlates with human judgment, and it interprets a higher hallucination score as a favorable sign, indicating the metric may capture something other than quality. A human-correlation check would settle whether the central claim holds; if it fails, the paper's conclusions about Nepali adaptation are unverified. The paper deserves credit for its organic Nepali pretraining data, transparent description of the two-stage DAPT procedure, and the attention-heatmap analysis, which is a reasonable qualitative probe. However, the disproportionate reliance on a single unvalidated metric for the main claim warrants keeping the CONDITIONAL verdict. My concern does not change the reader's verdict; it sharpens the condition: release the evaluation prompts, the model outputs, and human-rated samples, and demonstrate agreement with GPT-4o. Therefore verdict_should_be is UNCHANGED, but I disagree with the reader's identification of the weakest assumption.","tokens_in":10915,"tokens_out":5051,"duration_ms":46802,"concrete_test":"Sample 50 question-answer pairs from the Nepali traffic exam used in Section 5.3, have two native Nepali speakers rate the generated outputs on the same five 0-10 attributes (correctness, grammar, usability, hallucination, overall quality), and compute Spearman correlation and agreement with the GPT-4o scores reported in the paper. If the correlation is below approximately 0.6, or if human judges rate correctness substantially lower than GPT-4o, the central claim that the model 'learned Nepali' is not supported. Additionally, verify that the exact evaluation prompt given to GPT-4o does not include the reference answer or otherwise bias the scoring.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's evidence for Q1—that the adapted model has learned Nepali—rests entirely on GPT-4o scores for five subjective attributes (Section 5.3, Figure 1). No scoring prompt is provided, no human correlation or inter-rater reliability is reported, and the paper's own Limitations section admits that human evaluators, especially Nepali experts, would give a more definitive assessment. The base model frequently produces empty or degenerate outputs that are scored 0, so the comparison may largely reflect fluency gains rather than semantic correctness. This concern is sharpened by the paper's interpretation of higher hallucination scores as a positive sign (Section 6), which suggests the metric may reward verbosity or confidence rather than factual accuracy. The central claim 'the model has learned Nepali' depends on this unvalidated automatic judge: if GPT-4o systematically favors fluent but incorrect Nepali, the claim is unsupported. The reader's data-quality concern is real but secondary: in the 5M parallel pairs the Nepali side is organic (OSCAR text), so translationese risk mainly affects the 114K instruction triplets, not the main pretraining corpus. Thus the most load-bearing assumption is the validity of the GPT-4o evaluation, not the synthetic-data quality.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates domain-adaptive pretraining (DAPT) of Llama 3 8B for Nepali using 4-bit QLoRA and synthetic parallel data (about 5M Nepali–English pairs and 114K translated instruction triplets). The pipeline has two pretraining stages (English-to-Nepali translation and bilingual next-token prediction) followed by mixed-language instruction finetuning. The authors compare base and adapted models on three questions: Nepali generation quality (GPT-4o scores on 78 license-exam questions), English forgetting (MMLU, ARC, Winogrande, TruthfulQA), and dependency resolution (self-attention heatmaps for adjective–noun pairs). They report forgetting on English benchmarks, larger relative few-shot gains for the adapted model (up to 19.29% on ARC-Challenge), which they interpret as latent retention, and qualitatively different attention patterns for Nepali adjectives.","tokens_in":11160,"tokens_out":5315,"duration_ms":45081,"significance":"If the claims were fully supported, the paper would provide a useful low-resource recipe: adapting an 8B model to an unsupported language using only synthetic data and modest compute, with evidence about forgetting and knowledge transfer. The strengths are the practical setup (QLoRA, open weights), a transparent data-generation pipeline (NLLB and IndicTrans2 with chrF++ filtering), and evaluation on four standard English benchmarks. However, the headline claims ('learned Nepali', 'latent retention', 'dependency resolution') currently rest on an unvalidated automatic judge, unquantified visual analyses, and internally inconsistent tables. With added validation and statistical support, the paper could serve as a useful empirical case study for DAPT in low-resource languages.","major_comments":[{"comment":"The central evidence for Q1 is GPT-4o scoring of 78 Nepali answers, but no scoring prompt, no human correlation, and no inter-rater reliability are reported; the Limitations section concedes that human evaluators would give a more definitive assessment. Because empty/degenerate base-model outputs are scored 0, the comparison may largely reflect fluency rather than semantic correctness, and Section 6's interpretation of higher hallucination scores as positive ('content that is more verifiable') suggests the scoring is not validated as factual accuracy. Please provide the full scoring prompt, report per-item scores, validate against at least a small set of Nepali human judgments, and justify or remove the hallucination-as-positive interpretation.","section":"§5.3, §6, Figure 1"},{"comment":"There are material inconsistencies between the table and text: Table 1 reports final-model Winogrande 0.5801/0.6275, but the text reports 0.5691/0.6022; TruthfulQA MC1/MC2 are 0.2827/0.4351 in the table but 0.2607/0.4243 in the text. These are not typo-level differences for the forgetting claim. Please correct the numbers and ensure the abstract, table, and text all use the same final values.","section":"§6, Table 1"},{"comment":"The 'latent retention' interpretation is not supported by significance testing or error bars. Relative percent improvements over a much lower 0-shot baseline (e.g., ARC-Challenge: 0.3183→0.3797, +19.29%, vs. base 0.5017→0.5179, +3.23%) conflate baseline level and ceiling effects; the absolute gain is larger for the final model, but no variance or significance is reported. Please report multiple evaluation runs or standard errors, and ideally test the few-shot improvement against a control adaptation (e.g., Nepali-only continued pretraining) before claiming latent retention.","section":"§6, Table 1, Abstract"},{"comment":"The Q3 conclusion that the final model 'has learned to attend to Nepali adjectives the way the base model attends to English ones' is based on visual inspection of heatmaps, with no quantitative summary, no comparison to random or baseline attention, and no measure of inter-annotator agreement. Please provide a quantitative metric (e.g., mean adjective-to-noun attention strength per layer/head, with error bars) and a significance test or permutation baseline.","section":"§6, Figure 2"}],"minor_comments":[{"comment":"The title contains a typo: 'Domain-adaptative' should be 'Domain-adaptive'.","section":"Title"},{"comment":"The sentence 'suggesting that our model leverages few-shot examples more effectively than the final model' should read '...more effectively than the base model.'","section":"§6, paragraph on few-shot gains"},{"comment":"The TruthfulQA MC1 and MC2 rows have empty 5-shot columns; please clarify why 5-shot results are not reported for these benchmarks.","section":"Table 1"},{"comment":"No human verification of the translated instruction sets is reported; the chrF++ backtranslation filter is a useful proxy, but the paper should state how many instruction triplets were discarded and give examples of retained and discarded items.","section":"§3.1"},{"comment":"Training hyperparameters (learning rate, batch size, sequence length, optimizer, number of steps, hardware) are not reported in sufficient detail to replicate the experiments; the statement 'training settings are much the same' is too vague.","section":"§3.2"},{"comment":"The claim that max-pooling is more suitable than mean-pooling is not supported by reported quantitative results; please provide the comparison that motivated this choice.","section":"§5.2.1"},{"comment":"The heatmaps and boxplots lack axis labels and color-scale details; the score distributions in Figure 1 would be easier to interpret if the attributes were explicitly labeled on the boxplot axes.","section":"Figures 1 and 2"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is a useful empirical report but is not yet at the level of a journal publication; the numerical inconsistencies and unvalidated evaluation are the main barriers. The weaknesses are methodological rather than ethical: the cited prior work is appropriate, and the issues can be addressed with additional validation, corrected tables, and error bars."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read of Duwal et al. The paper is a straightforward application of known DAPT + QLoRA + synthetic backtranslation methods to Nepali, and it does that competently. The training setup is sensible: parallel English–Nepali data where the Nepali side is organic OSCAR text, a translation pretraining step, a bilingual NTP step, then SFT on filtered instruction data. They compare NLLB and IndicTrans2, apply a chrF++ filter, and are transparent about compute limits (4-bit, single epoch). That is credit where it's due: it's a reproducible recipe, and it will be useful to people doing similar low-resource adaptation.\n\nThe soft spots are real. The central claim \"the model has learned Nepali\" rests entirely on GPT-4o scores for 78 traffic-license questions, with no scoring prompt, no human correlation, no inter-rater reliability, and no significance testing. The paper's own Limitations section concedes human experts would give a more definitive assessment. Given that the base model often produces empty outputs scored 0, the comparison may be mostly capturing fluency gains rather than semantic correctness. The interpretation of higher hallucination scores as a positive sign (Section 6) is not convincing—it reads as post-hoc spin, and it undercuts confidence in the judge's reliability. The \"latent retention\" claim from few-shot percent improvements is overinterpreted too: the adapted model's 0-shot scores are much lower, so percent gains are not a fair comparison, and no significance testing is reported. The attention heatmap analysis is qualitative, based on 26 Nepali adjective–noun pairs, with no quantitative summary.\n\nThere are also internal numerical inconsistencies: Winogrande and TruthfulQA numbers in the text differ from Table 1. That is minor but should be fixed. No code or data is actually released despite the abstract promising it.\n\nI agree with the stress-test note: the most load-bearing assumption is the validity of the GPT-4o evaluation, not the synthetic-data quality, since the main pretraining target is organic Nepali. The synthetic instruction data is a secondary concern.\n\nWho is this for? Practitioners working on low-resource language adaptation who want a concrete recipe and a baseline. It is not a methodological advance. As a paper, it deserves a serious referee—the question of whether synthetic data can adapt a model to a low-resource language is worth engaging with, and the recipe is executed cleanly. But the evaluation needs to be strengthened before publication. My recommendation: send it to peer review, with the expectation of major revision—require code/data release, human or validated automatic evaluation, correction of the inconsistencies, and a toned-down interpretation of the few-shot and attention results.","headline":"A clean low-resource adaptation recipe for Nepali, but the central claim rests on an unvalidated GPT-4o evaluator and some loose interpretation.","tokens_in":11695,"tokens_out":2550,"would_cite":false,"duration_ms":23452,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that domain-adaptive continual learning on synthetic Nepali–English parallel data, run in 4-bit QLoRA, adapts Llama 3 8B to generate Nepali and retain latent English knowledge.","keywords":["domain-adaptive pretraining","continual learning","low-resource languages","Nepali language","synthetic parallel data","QLoRA","catastrophic forgetting","attention heatmaps"],"falsifier":"A native-speaker evaluation of the adapted model's Nepali outputs would settle it: ask Nepali speakers to rate grammaticality and overall quality on the same 78 questions and compare with the base model. If human ratings do not reproduce the GPT-4o advantage, the central claim fails. A second check is to back-translate a sample of the 5M parallel pairs and have a bilingual annotator judge the Nepali side; a high rate of MT artifacts or unidiomatic constructions would undermine the data-generation premise.","tokens_in":10722,"feed_emoji":"🇳🇵","tokens_out":8232,"duration_ms":67982,"temperature":0.7,"pith_summary":"The paper claims that a large language model can be adapted to a low-resource language using only synthetic parallel data and a small fraction of trainable parameters. It continually trains Llama 3 8B on 5 million Nepali–English paragraph pairs and translated instruction data with 4-bit QLoRA, then compares the adapted model against the base model. The adapted model produces Nepali answers that GPT-4o scores substantially higher, especially on grammatical correctness, while English benchmarks show expected forgetting but larger few-shot improvements (up to 19.29% versus 4.98%), which the authors read as latent retention of English knowledge. If these results hold, domain-adaptive continual learning is a practical path for resource-constrained languages that lack native corpora or dedicated benchmarks.","feed_headline":"Synthetic data alone adapts Llama 3 to Nepali","feed_subtitle":"4-bit continued training lifts Nepali generation scores and boosts few-shot English gains to 19.29%.","key_machinery":"The central mechanism is a two-stage QLoRA continual pretraining on synthetic Nepali–English parallel data: first the model is trained to translate English to Nepali so it generates organic Nepali, then it is trained on bilingual next-token prediction with alternating Nepali and English sentences to reuse English knowledge. Both stages use rank-128 4-bit QLoRA, updating roughly 335M of Llama 3 8B's parameters, followed by a rank-16 instruction finetune on translated Alpaca, Dolly, and WebGLM sets. The parallel corpus is built with NLLB (8-bit) for 5M paragraph pairs and IndicTrans2 for 114K instruction triplets, filtered by a chrF++ backtranslation threshold of 50. For linguistic probing, the paper max-pools token-level self-attention into word-level attention and plots layer-head heatmaps for adjective–noun dependency pairs.","core_discovery":"On its own terms, the paper establishes that domain-adaptive pretraining on synthetic data moves a base model from almost no usable Nepali generation to recognizable, higher-scoring Nepali output. Using English-to-Nepali translation pretraining followed by bilingual next-token prediction, the final model's answers to 78 Nepali traffic-license questions score higher than the base model across correctness, grammar, usability, hallucination, and overall quality, with grammar showing the clearest improvement. On English benchmarks the adapted model forgets, but the relative gain from 0-shot to 5-shot prompting is consistently larger than the base model's, reaching 19.29% on ARC-Challenge, a pattern the paper interprets as evidence that English knowledge is retained latently and can be reactivated through in-context examples. The attention heatmaps finally show the adapted model attending from Nepali adjectives to their governing nouns in a way that resembles the base model's English patterns, supporting the claim that the model has acquired structural knowledge of Nepali rather than just token-level mimicry.","pith_inferences":["The 'latent retention' reading could be tested directly by building Nepali versions of ARC or MMLU: if few-shot gains come from retained knowledge, the adapted model should also improve on Nepali-language reasoning tasks, which the paper notes do not yet exist.","A native-speaker rating study of the same 78 generated answers would separate genuine Nepali competence from GPT-4o's tolerance for translationese; the paper itself names human evaluation as its main missing check.","Because the synthetic corpus is dominated by online news text, the adapted model's grammatical gains may be register-specific; a test on conversational or dialectal Nepali would show whether the dependency-resolution patterns generalize.","The method's parameter efficiency suggests a sequential-adaptation possibility the paper leaves implicit: one base model could be adapted to several low-resource languages in turn, with per-language LoRA adapters, and the few-shot reactivation effect could serve as shared retrieval across those languages."],"forward_implications":["A low-resource language without an instruction corpus can inherit English knowledge through a translation-and-bilingual-pretraining loop, so Nepali-style adaptation may be repeatable for other South Asian languages.","Catastrophic forgetting after domain-adaptive pretraining is real but not permanent in practice: the adapted model recovers more of its English benchmark performance when given a few examples, so few-shot prompting should be part of the evaluation protocol for continually trained models.","Attention heatmaps of adjective-noun relations offer a cheap way to check whether an adapted model has learned syntax rather than surface token statistics."],"supporting_citations":[{"why":"Introduces domain-adaptive pretraining (DAPT) and its gains in low-resource settings; the paper's method is built on this paradigm.","marker":"(Gururangan et al., 2020)"},{"why":"Supplies the two-stage recipe of translation pretraining plus bilingual next-token prediction, which the paper adapts to Nepali.","marker":"(SarvamAI, 2023)"},{"why":"Provides QLoRA, the 4-bit low-rank training method used for both continual pretraining and finetuning.","marker":"(Dettmers et al., 2023)"},{"why":"NLLB, chosen as the paragraph translator after FLORES evaluation, generates the Nepali halves of the 5M parallel pairs.","marker":"(Costa-jussà et al., 2022)"},{"why":"IndicTrans2, used to translate English instruction sets into Nepali for the finetuning step.","marker":"(Gala et al., 2023)"},{"why":"FLORES Nepali–English test set, used to choose NLLB over the alternative translation systems.","marker":"(Guzmán et al., 2019)"},{"why":"Documents GPT-4/GPT-4o multilingual generation strengths, which the paper relies on for automatic scoring of Nepali answers.","marker":"(OpenAI, 2024)"},{"why":"Shows that attention heads encode syntactic relations; the paper's adjective-noun heatmap analysis follows this approach.","marker":"(Vig and Belinkov, 2019)"}],"fun_headline_variants":["Synthetic data adapts Llama 3 to Nepali, with latent English retention","Low-resource Nepali adaptation: synthetic data boosts few-shot gains","Adapting Llama 3 to Nepali with synthetic data finds latent retention","Synthetic data for Nepali continual learning reveals structural knowledge","Continual learning on synthetic Nepali data: Llama 3 retains English"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the synthetic Nepali side of the parallel corpus is genuinely valid, representative Nepali: the paper reports no human quality check on the 5 million translated paragraph pairs, so if those translations are translationese or contain systematic machine-translation errors, every downstream claim about the model learning Nepali loses its foundation.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic data adapts Llama 3 to Nepali, with latent English retention","Low-resource Nepali adaptation: synthetic data boosts few-shot gains","Adapting Llama 3 to Nepali with synthetic data finds latent retention","Synthetic data for Nepali continual learning reveals structural knowledge","Continual learning on synthetic Nepali data: Llama 3 retains English"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000528,"raw_usage":{"total_tokens":2571,"prompt_tokens":991,"completion_tokens":1580,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1483}},"tokens_in":607,"tokens_out":1580,"duration_ms":10775,"temperature":1.0,"reasoning_tokens":1483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:42:37.445041+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A native-speaker evaluation of the adapted model's Nepali outputs would settle it: ask Nepali speakers to rate grammaticality and overall quality on the same 78 questions and compare with the base model. If human ratings do not reproduce the GPT-4o advantage, the central claim fails. A second check is to back-translate a sample of the 5M parallel pairs and have a bilingual annotator judge the Nepali side; a high rate of MT artifacts or unidiomatic constructions would undermine the data-generation premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the two-stage recipe of translation pretraining plus bilingual next-token prediction, which the paper adapts to Nepali."}],"review_version":1}