{"id":"cf5ec44e-c776-48c0-a5d2-541f29af758b","arxiv_id":"2605.27379","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.","lead":"Soro adapts open Gemma 3 models into Tajik chatbots via continual pretraining and instruction tuning, then quantizes them for school deployment. It ships new Tajik benchmarks and reports large gains over same-size baselines in a live 100-school pilot.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"School-aligned benchmarks may partly measure memorization of the same textbook class used in continual pretraining, so reported Tajik gains are not fully independent of train data.","rationale":"The reader correctly isolates the primary scientific risk: train–test proximity between Ministry textbooks in §4.1 and the grade 5–11 History/Literature benchmarks in §A.4/§A.6. That risk is load-bearing for the claim of “substantial” generalizable Tajik gains, because those sets contribute to the multi-benchmark average and show large lifts. FactQA (entrance exams) and TajLib are more independent and still show gains, so the paper is not empty; contamination would force a narrower claim rather than collapse the whole result. Synthetic SFT quality is real but secondary. No stronger internal inconsistency appears: methods are standard LoRA CPT + SFT + linear merge; English MMLU retention and quantization tables are consistent with the deployment story. Therefore the reader’s CONDITIONAL verdict (accept-shaped if contamination is audited and artifacts released) should stand; the concrete decontamination check is the single experiment that settles whether the concern lands. Agreement with the reader is full on the weakest assumption.","tokens_in":35012,"tokens_out":694,"duration_ms":7098,"concrete_test":"Run 13-gram (or 8–13-gram) and embedding-based decontamination of Tajik History and Tajik Literature against the educational textbook subset of the 1.9B corpus (or a released sample of it). Report contamination rate and recompute Fig. 4 / per-benchmark accuracies after removing contaminated items (or after a held-out textbook split). If cleaned accuracy drops by more than ~3 points relative to Gemma 3-IT, the independence assumption fails and the strongest claim must be narrowed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that Tajik-only continual pretraining on a 1.9B-token corpus (including curriculum-aligned educational materials) plus SFT and an ~80/20 merge yields substantial, generalizable gains on Tajik benchmarks over same-size Gemma 3-IT (e.g., Fig. 4 averages 58.2 vs 50.4 for 12B; 63.8 vs 55.8 for 27B). The least secure premise is independence of the school-aligned eval sets. §4.1 states the pretraining corpus includes secondary-school materials grades 5–11 (textbooks, manuals, lecture notes) from Ministry of Education sources via OCR/transcription. Appendices A.4 and A.6 state that Tajik History (1,400 items) and Tajik Literature (698 items) are derived from the same class of officially approved school textbooks for grades 5–11. No n-gram, embedding, or passage-level decontamination between this educational subset and those benchmarks is reported. If substantial overlap exists, part of the ~6–8+ point average lift (and larger lifts on History/Literature/TajLib) can be explained by memorization of training material rather than improved general Tajik competence. Gains on Tajik-FactQA (university entrance exams) and TajLib reduce but do not eliminate the concern for the headline multi-benchmark average. Synthetic Gemini SFT is secondary.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces Soro (27B) and Soro Lite (12B), Tajik-specialized conversational LLMs obtained from Gemma 3 via Tajik-only continual pretraining on a curated 1.9B-token corpus (web, PDFs, curriculum materials), supervised instruction tuning on 40K teacher-style examples, and an ~80/20 linear merge with Gemma 3-IT. The authors release a suite of Tajik multiple-choice benchmarks (Tajik MMLU, Tajik-FactQA, History, Literature, TajLib, Curated), report average accuracy gains of roughly 6–8+ points over same-size Gemma 3-IT while largely retaining English MMLU, show that FP8 and GPTQ INT4 preserve most Tajik gains, and describe an education-sector pilot across 100 schools with a teacher feedback survey.","tokens_in":35440,"tokens_out":1348,"duration_ms":22699,"significance":"If the reported Tajik gains are largely free of train–eval contamination and statistically reliable, this is a concrete, reproducible contribution for a severely under-resourced language: open Tajik benchmarks on Hugging Face, a documented continual-pretraining + SFT + merge recipe under modest compute, quantization results that enable edge deployment, and a government-backed school pilot with quantitative teacher ratings. The tokenizer fertility analysis motivating Gemma 3, the merge-ratio sweep (Fig. 3), and the pre-merge / post-merge / quantized ablations are useful methodological details for other low-resource adaptation efforts. The work is more systems-and-deployment than algorithmic novelty, but that is appropriate for the stated goal of deployable Tajik language technology.","major_comments":[{"comment":"Independence of school-aligned benchmarks from the pretraining corpus is not established and is load-bearing for the headline claim of generalizable Tajik gains. §4.1 states that continual pretraining includes secondary-school materials (grades 5–11 textbooks, manuals, lecture notes) from Ministry sources via OCR/transcription. Appendices A.4 and A.6 state that Tajik History (1,400 items) and Tajik Literature (698 items) are derived from the same class of officially approved grades 5–11 textbooks. No n-gram, embedding, or passage-level decontamination between the educational subset and these benchmarks is reported. Part of the multi-benchmark average lift in Fig. 4 (and larger lifts on History/Literature) may therefore reflect memorization. Please report decontamination statistics and either (i) re-evaluate after removing overlapping items or (ii) report headline averages excluding Histo","section":null},{"comment":"Statistical support for accuracy claims is thin. Fig. 4, Table 6, and the per-benchmark figures report point accuracies only—no bootstrap CIs, standard errors, or significance tests on Soro vs Gemma 3-IT (or vs other open models). With fixed multiple-choice sets and deterministic logit extraction (Appendix B), resampling over items is straightforward. Without uncertainty estimates, statements such as “substantially outperforms” and “preserves most Tajik-language gains” under quantization cannot be assessed for robustness, especially on smaller sets (e.g., Tajik-FactQA, n=436).","section":null},{"comment":"Synthetic data provenance and quality controls are under-specified relative to their role in both pretraining and SFT. FineWeb-Edu is translated into Tajik with Gemini 2.5 Flash (§4.1), and the 40K instruction set is also Gemini-generated with a human audit of only “several hundred” examples and a qualitative “low defect rate” (§4.2)—no defect-rate numbers, inter-annotator agreement, or subject-stratified failure modes. Because the chatbot’s pedagogical style and much of the educational content rest on this pipeline, the manuscript should quantify audit outcomes and discuss residual risks (hallucinated facts, English interference, style homogenization) as limitations on the educational claims.","section":null}],"minor_comments":[{"comment":"Abstract and §6 claim average gains on the order of 6–8+ points and “8–13%” in the conclusion; align the wording (percentage points vs percent) and state explicitly which model variants and which benchmark average are meant.","section":null},{"comment":"Fig. 4 and related bar charts would be clearer with error bars once item-level uncertainty is computed; also label whether scores are raw-logit or chat-template format (Appendix B describes both).","section":null},{"comment":"Table 2 fertility comparison is useful; briefly note whether fertility was measured with the same normalization (whitespace, punctuation) for Cyrillic Tajik vs English so readers can interpret the 2.380 vs 2.798 gap.","section":null},{"comment":"§5.2 merge experiment: state whether the 80/20 optimum was chosen on a held-out split of the Tajik suite or on the full evaluation set used for final reporting, to avoid selection bias on the merge weight.","section":null},{"comment":"Appendix E teacher survey (n=53, convenience sample of pilot schools) is a welcome addition; the limitations paragraph already notes selection bias—consider moving a one-sentence caveat into the main §8.4 so readers of the deployment case study see it without opening the appendix.","section":null},{"comment":"Minor consistency: date line “April 2026” / arXiv May 2026 and survey dates “March 16–30, 2026” are fine if intentional; ensure all URLs and HF dataset names remain stable at camera-ready.","section":null}],"recommendation":"major_revision","confidential_remarks":"The contamination concern is the main reason I recommend major_revision rather than minor_revision: it is fixable with analysis and clearer reporting, not necessarily with retraining, but until it is addressed the central scientific claim is hard to credit at face value. Novelty is primarily in the language/resource setting and open evaluation suite rather than new methods; that is still a good fit for a systems-oriented AI venue if the evaluation is cleaned up. I would not reject on novelty grounds alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean applied multilingual paper: Gemma 3 12B/27B, Tajik-only continual pretraining on 1.9B tokens (FineWeb2, FinePDFs, translated FineWeb-Edu, plus Ministry textbooks via OCR), LoRA, 40K Gemini teacher-style SFT with a small native audit, then ~80/20 linear merge with Gemma-3-IT. They ship Soro / Soro Lite, open a six-set Tajik MC suite on HF, show ~6–8+ point average lifts over same-size Gemma-3-IT (Fig. 4: 58.2 vs 50.4; 63.8 vs 55.8) with English MMLU largely held, and FP8/INT4 that keep most of the Tajik gain for edge boxes. The 100-school pilot under NAIS-2040 is unusually concrete for this literature.\n\nWhat is new is not the recipe—continual pretrain, LoRA, SFT, linear merge, GPTQ are standard and cited—but the Tajik-centric artifacts, curriculum-aligned data mix, open eval suite (FactQA, TajLib, History, Literature, Curated, Tajik MMLU), fertility analysis that actually motivates Gemma 3, and documented deployment with teacher survey numbers. Citation pattern is appropriate; no load-bearing math to break.\n\nThe stress-test concern is real but partial. History and Literature are drawn from the same class of grades 5–11 Ministry textbooks that sit in the pretraining mix, and the paper reports no n-gram or passage decontamination. So part of the lift on those two sets (and the multi-benchmark average) could be memorization. FactQA (entrance exams) and TajLib (linguistic competence) still move, which keeps the central claim from collapsing. Secondary soft spots: no error bars or significance tests, and heavy reliance on Gemini for SFT with only a few-hundred-item human audit. Those are fixable with an audit appendix and uncertainty estimates, not reasons to dismiss the work.\n\nWho it is for: people building low-resource language models, EdTech under connectivity constraints, and anyone who wants a transferable blueprint rather than a new principle. I would bring it to reading group as a systems/deployment case, cite the benchmarks and the pilot numbers if I work on similar languages, and send it to peer review as an applications contribution—conditional on contamination checks and full artifact release, not desk reject.","headline":"Solid applied Tajik specialization with open benchmarks and a real 100-school pilot; the main soft spot is possible textbook train–test overlap, not empty claims.","tokens_in":36083,"tokens_out":614,"would_cite":true,"duration_ms":7279,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Tajik-only continual pretraining on 1.9B tokens turns open Gemma 3 into deployable Soro chatbots that beat same-size baselines on new Tajik school and language tests, and still run after FP8/INT4 compression for schools with weak connectivi","keywords":["Tajik language models","continual pretraining","low-resource NLP","instruction tuning","model merging","quantization","educational AI deployment","language-specific benchmarks"],"falsifier":"A decontamination or overlap audit between the released Tajik History and Literature benchmarks and the educational subset of the 1.9B-token pretraining corpus: if many questions or near-paraphrases already appear in training, the claimed generalization gains shrink; independent human evaluation on held-out local topics would also test whether synthetic instruction data introduced systematic bias.","tokens_in":35931,"feed_emoji":"🇹🇯","tokens_out":1153,"duration_ms":15091,"temperature":0.7,"pith_summary":"This paper claims that a focused, Tajik-only adaptation of open Gemma 3 models is enough to produce practical conversational systems for a language that mainstream multilingual training largely ignores. The authors start from 12B and 27B Gemma 3 checkpoints, continually pretrain on a curated 1.9-billion-token Tajik corpus that mixes filtered web text, PDFs, and Ministry-aligned school materials, then instruction-tune on 40K teacher-style Tajik examples and linearly merge back about 20 percent of the original Gemma instruction weights. They also release a new suite of Tajik multiple-choice benchmarks covering general knowledge, linguistic competence, national history and literature, and university-entrance material. On those benchmarks the final Soro models substantially outperform same-size Gemma 3 instruction-tuned baselines while largely holding English MMLU, and FP8 and INT4 quantization keep most of the Tajik gains while cutting memory enough for edge or single-GPU use. The work is already running in a government- and UNICEF-linked pilot across 100 schools in five Tajik cities, with a stated path toward national scale-out under the country’s AI strategy.","feed_headline":"Tajik-only training turns Gemma 3 into school-ready Soro","feed_subtitle":"1.9B tokens, new Tajik exams, and FP8/INT4 keep the gains small enough for edge classrooms","key_machinery":"The three-stage pipeline of Tajik-only LoRA continual pretraining on a 1.9B-token curated corpus, supervised instruction tuning on 40K teacher-style examples, and linear weight merging (~80% Soro / ~20% Gemma 3-IT), together with the open-sourced Tajik evaluation suite that makes the gains measurable.","core_discovery":"After Tajik-only continual pretraining, instruction tuning, and an approximately 80/20 linear merge with the original Gemma 3 instruction-tuned weights, the resulting Soro (27B) and Soro Lite (12B) models achieve clear average gains over same-size Gemma 3-IT baselines on a new suite of Tajik benchmarks while retaining most English capability; FP8 and GPTQ INT4 quantization preserve the bulk of those Tajik gains at far lower memory cost, making school and edge deployment feasible.","pith_inferences":["If textbook-to-benchmark contamination is low, the same recipe is a practical template for other Central Asian and Persian-related languages that share Cyrillic or Arabic-script under-representation.","The large gap on the pure linguistic competence set (TajLib) suggests that continual pretraining mainly improves internal grammar and morphology representations, not only surface fact recall.","Teacher survey scores that rate language quality highest and factual accuracy lower imply that future preference or RAG layers on verified local knowledge bases would address the remaining deployment friction more than further pretraining alone.","The edge-quantization story makes the model transferable to other low-connectivity public-sector settings beyond education, provided local knowledge bases and safety filters are added."],"forward_implications":["Low-resource languages with similar sparse digital footprints can obtain large, measurable gains by continual pretraining a strong open multilingual base on a carefully curated monolingual corpus rather than training from scratch.","Linear merging with a modest fraction of the original instruction-tuned weights can recover general English and world knowledge lost during language-specific adaptation without erasing the new language gains.","FP8 and INT4 quantization that preserve most specialized-language accuracy enable single-GPU and consumer-GPU deployment of 12B–27B models in schools and offices that lack data-center hardware or reliable connectivity.","Open-sourcing language-specific school and entrance-exam benchmarks fills an evaluation vacuum and lets others measure progress on the same culturally grounded tasks.","A curriculum-aligned training corpus plus teacher-style instruction data supports real classroom use cases (lesson help, diagnostics, AI literacy) already being piloted at national scale."],"fun_headline_variants":["Tajik-only pretraining turns Gemma 3 into edge-ready Soro","Soro: Gemma 3 specialized for Tajik schools via 1.9B tokens","Continual Tajik training yields Soro that beats Gemma 3 on local exams","FP8/INT4 Soro keeps Tajik gains small enough for classrooms","New Tajik benchmarks show Soro outperforming same-size Gemma 3"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That the new school-aligned Tajik benchmarks are sufficiently independent of the same class of Ministry textbooks that were transcribed or OCR’d into the pretraining corpus, so reported gains reflect general competence rather than memorization of training material.","fun_headline_variants_meta":{"raw":{"variants":["Tajik-only pretraining turns Gemma 3 into edge-ready Soro","Soro: Gemma 3 specialized for Tajik schools via 1.9B tokens","Continual Tajik training yields Soro that beats Gemma 3 on local exams","FP8/INT4 Soro keeps Tajik gains small enough for classrooms","New Tajik benchmarks show Soro outperforming same-size Gemma 3"]},"model":"grok-4.5","effort":"low","cost_usd":0.006208,"raw_usage":{"total_tokens":1622,"prompt_tokens":785,"num_sources_used":0,"completion_tokens":108,"cost_in_usd_ticks":62080000,"prompt_tokens_details":{"text_tokens":785,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":729,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":785,"tokens_out":108,"duration_ms":6827,"temperature":1.0,"reasoning_tokens":729,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T23:56:32.629934+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A decontamination or overlap audit between the released Tajik History and Literature benchmarks and the educational subset of the 1.9B-token pretraining corpus: if many questions or near-paraphrases already appear in training, the claimed generalization gains shrink; independent human evaluation on held-out local topics would also test whether synthetic instruction data introduced systematic bias.","supporting_citations":[],"review_version":1}