{"id":"db919f77-4559-4df4-83fc-723df45c3e01","arxiv_id":"2412.16423","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fine-tuned 1.2B Japanese medical SLM tops 6 of 8 JMED-LLM tasks against larger models, but the comparison is confounded by benchmark-specific fine-tuning.","lead":"A 1.2 billion parameter Japanese language model for clinical and medical text was trained, fine-tuned, and scored highest on 6 of 8 JMED-LLM tasks compared with much larger models. The result is weakened because the model was fine-tuned on the benchmark's own training data, and some baseline scores were not evaluated under the same protocol.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '6 of 8 best' JMED-LLM claim compares a model fine-tuned on the benchmark's training split with external baselines that were not fine-tuned; this confounds performance with benchmark exposure.","rationale":"The reader's weakest-assumption analysis identifies the same core issue: the comparison is only meaningful if other models were evaluated under the same protocol and if it is fair to compare a model fine-tuned on JMED-LLM training data with models that were not. This is the single most load-bearing concern because it directly determines whether the headline claim supports the paper's broader conclusion. If the matched-fine-tuning check were run and showed comparable performance by an 8B baseline, the '6 of 8 highest' result would be an artifact of benchmark-specific training rather than an SLM advantage. The paper is otherwise transparent: training setup, tokenizer, ablations, and limitations are all described, and the author honestly reports that base models fail on JMED-LLM and that synthetic exercises sometimes hurt IgakuQA performance. Those strengths do not remove the need for a matched comparison. The reader's verdict of CONDITIONAL is appropriate, so no verdict change is needed.","tokens_in":14837,"tokens_out":3216,"duration_ms":27875,"concrete_test":"Instruction-tune one strong baseline (e.g., Llama-3-ELYZA-JP-8B-Instruct or Qwen2-7B-Instruct) on the exact same JMED-LLM training split (all tasks except test samples) using the same instruction-tuning format and comparable hyperparameters, then evaluate on the same JMED-LLM test split. If the tuned baseline matches or exceeds NCVC-slm-1-instruct on the 6 tasks, the central 'SLM outperforms larger models' claim fails. If the tuned baseline remains below NCVC-slm-1-instruct on most tasks, the small-model advantage is supported. Also publish the evaluation prompts and decoding settings to verify identical protocol with the external baseline scores.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 3.3 is that NCVC-slm-1-instruct achieved the highest scores on 6 of 8 JMED-LLM tasks compared with larger models including GPT-4o. However, Section 2.4.2 states that instruction tuning 'was used 8 JMED-LLM dataset except test samples,' so the model was trained on the same task distribution as the evaluation set. The comparison models' scores are cited from external sources [67] and were not fine-tuned on those JMED-LLM training sets. The base NCVC-slm-1 models score near zero on these tasks (Tables 3 and 4); after fine-tuning, they jump to 0.98 kappa on SMDIS and 0.88-0.90 partial F1 on NER tasks. This indicates that the improvement is largely attributable to supervised training on the benchmark's training split, not to a general small-model advantage. As published, the headline result does not establish that an SLM outperforms larger models; it establishes that a model fine-tuned on the evaluation benchmark's training data scores higher than models evaluated zero-shot. That is an experimentally real but much weaker claim, and it is the load-bearing condition for the paper's conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the development of NCVC-slm-1, a 1.2B-parameter Japanese language model for clinical and medical text, trained on filtered Wikipedia and OSCAR data supplemented with scraped and synthetic medical textbooks. After instruction tuning on the JMED-LLM training data, the model is evaluated on IgakuQA and JMED-LLM; the authors report that the instruction-tuned model achieves the highest scores on 6 of 8 JMED-LLM tasks compared with larger models including GPT-4o. The paper also reports ablations on training token budgets and on the inclusion of synthetic exercises, and it discusses design choices for tokenization, architecture, and training stability.","tokens_in":15049,"tokens_out":3842,"duration_ms":33821,"significance":"If the comparative claims were properly supported, the work would be a practically useful demonstration that a small, locally deployable model can compete with much larger models on several Japanese clinical and medical NLP tasks after modest fine-tuning. The paper is transparent about many engineering details, including data filtering results, tokenizer design, architecture decisions, and negative results on IgakuQA and on synthetic-data fine-tuning. However, the central '6 of 8 best' claim is currently confounded: the proposed model was instruction-tuned on the JMED-LLM training splits while the comparison models were evaluated zero-shot with scores taken from an external source. The engineering contribution is real, but the headline comparative result needs to be reframed or re-evaluated before it can be accepted.","major_comments":[{"comment":"The central claim that NCVC-slm-1-instruct 'was the highest score of 6 tasks' is confounded with benchmark exposure. Section 2.4.2 states that the instruction tuning dataset used '8 JMED-LLM dataset except test samples,' so the model was trained on the same task distribution as the evaluation sets. The comparison models' scores in Table 3 and Table 4 are cited from reference [67], and the paper gives no indication that those baselines were fine-tuned on JMED-LLM training data. This is not a minor caveat: the base NCVC-slm-1 models score near zero on most JMED-LLM tasks (e.g., 0.00 F1 on all NER tasks in Table 4), and after fine-tuning they jump to 0.98 kappa on SMDIS and 0.87-0.90 partial F1 on NER tasks. The jump indicates that the results are largely attributable to supervised training on the benchmark itself. To support the comparative conclusion, the authors should either (a) fine-tune the comparison models on the same JMED-LLM training data, (b) evaluate NCVC-slm-1 under the same zero-shot protocol used for the baselines, or (c) explicitly and prominently reframe the claim as demonstrating that fine-tuning an SLM on a target benchmark yields high scores relative to zero-shot baselines.","section":"Section 2.4.2 / Section 3.3 / Tables 3-4"},{"comment":"The pre-training corpus contains synthesized exercises described as 'similar to Japanese national medical licensing examinations,' and the paper evaluates on IgakuQA, which is exactly such examinations. No overlap or contamination analysis is reported between the synthetic exercises and the IgakuQA test items (or the JMED-LLM test items). Given that the synthetic exercises were generated from disease and drug topic lists, it is plausible that near-duplicates of real exam questions appear in the training data. The authors should provide a quantitative overlap analysis (for example, n-gram overlap between the synthetic exercise set and the evaluation sets) and discuss the implications for the reported IgakuQA and JMED-LLM numbers. This is particularly relevant for interpreting the ablation in Table 5, where adding synthetic exercises degrades IgakuQA performance.","section":"Section 2.1.2 / Section 2.5 / Table 2"},{"comment":"No confidence intervals, error bars, or significance tests are reported for any JMED-LLM score, and several differences between NCVC-slm-1-instruct and the best baseline are small (for example, MRNER-medicine is 0.65 for NCVC-slm-1-instruct-20B and 0.65 for gemma-2-9b-it; JCSTS is 0.75 versus 0.60). Because the baseline numbers come from an external, unofficial source ([67], a speaker-deck), the score differences may not be reproducible under a single evaluation harness. The authors should provide their evaluation code and run the comparison models under the same protocol, or at minimum report the seed-averaged variance of the scores, before claiming superiority on 6 of 8 tasks.","section":"Tables 3-4 / Section 3.3"}],"minor_comments":[{"comment":"The abstract says '1B parameters' while Section 2.3 states 'approximately 1 billions (more accurately 1.2B)'; please use one consistent number.","section":"Abstract / Section 2.3"},{"comment":"The text lists 'MRNER-disease, MRNER-disease, and NRNER'; the second instance should be MRNER-medicine.","section":"Section 3.3 / Table 4 legend"},{"comment":"The row label 'Llama-3-youko-8b-insturct' contains a typo (insturct instead of instruct).","section":"Table 3"},{"comment":"The caption for Table 2 should specify that the NCVC-slm-1-instruct rows are tuned with JMED-LLM only and without the synthetic exercises dataset; the current text reports these rows but the comparison with Table 5 is not immediately clear.","section":"Section 3.2 / Table 2"},{"comment":"The sentence 'Their textbooks were 5 versions of previous language models' is grammatically unclear; it should read 'These textbooks were generated by 5 versions of previous language models.'","section":"Section 2.1.2"},{"comment":"The conclusion contains a duplicated word: 'This type of SLM is expected to to assist human jobs.'","section":"Section 5"},{"comment":"Figure 6 shows cyclic deterioration and improvement of base-model IgakuQA scores with increasing pretraining tokens, but the main text does not discuss this pattern or its possible causes.","section":"Appendix A.4"}],"recommendation":"major_revision","confidential_remarks":"This is a technical report with a strong engineering component, but the headline comparison is not yet supported because of the training-set confound and the reliance on external, unofficial baseline scores. The paper's own text in Sections 2.4.2 and 2.1.2 provides the evidence for this concern, so a revision that reframes the claims and adds proper controls would be feasible. Given the paper's current framing as a comparative benchmark result, I would not recommend acceptance without addressing the three major comments. The reference [67] is a speaker-deck; the authors should either cite the underlying dataset paper or reproduce those scores with their own harness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the model and the training write-up are useful, but the 'beats GPT-4o on 6 of 8 JMED-LLM tasks' claim does not survive contact with Section 2.4.2. The instruct models were fine-tuned on the JMED-LLM training splits; the baselines are zero-shot scores taken from a blog post and a slide deck. That is an unfair comparison, and the stress-test note correctly flags it as load-bearing.\n\nWhat is actually new: NCVC-slm-1 is a new 1.2B Japanese clinical model, with a careful morphological-analyzer tokenizer, a J-Medic-based dictionary, and synthetic textbook augmentation following the phi-1 recipe. The report documents hyperparameters, training curves, and honest ablations. I give real credit for the negative results: the LLM-based quality filter barely moved the data, and adding synthetic exercises to instruction tuning often hurt IgakuQA performance. That kind of candid reporting is rare and useful.\n\nThe soft spots are real but bounded. The main one is the evaluation protocol. The base model scores near zero on JMED-LLM and jumps to 0.98 kappa after fine-tuning on the benchmark's own training data. That says supervised adaptation works, which is not the same as a small model outperforming GPT-4o. The external baseline numbers [67] are not audited, and there are no confidence intervals or significance tests. A second issue: the pretraining corpus includes synthetic exercises similar to the Japanese national medical exam, and IgakuQA is built from those exams. The author acknowledges the risk but does not test for overlap. The report is a technical report, not a peer-reviewed paper; there is also no code or model release, which limits reproducibility.\n\nNone of this sinks the paper's more measured message: a 1B model, fine-tuned on a specific clinical benchmark, can be effective on several tasks and cheap enough to run locally. That is plausible from the tables, given the caveat about benchmark-specific training.\n\nBottom line: this is a paper for people building Japanese clinical NLP systems or studying SLM efficiency. It deserves a serious referee, mainly to force a fair baseline comparison and decontamination check. I would bring it to a reading group as a cautionary case study in benchmark evaluation. I would not cite the headline claim, but I would cite the model and training details.","headline":"A useful, honest Japanese clinical SLM report whose headline 'beats GPT-4o' is an artifact of fine-tuning on the benchmark's own training split.","tokens_in":15601,"tokens_out":3003,"would_cite":true,"duration_ms":27446,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 1.2B-parameter Japanese clinical model claims top scores on six of eight medical NLP tasks and runs on about 2.2 GB of GPU memory.","keywords":["small language model","Japanese clinical NLP","instruction tuning","medical language model","JMED-LLM","morphological analysis","synthetic textbooks","named entity recognition"],"falsifier":"Run the same eight JMED-LLM tasks under one protocol for all models, using a version of NCVC-slm-1 trained only on the allowed training split and comparison models evaluated in the same harness; if the small model no longer tops six tasks, the central claim is falsified. Separately, search for n-gram or embedding-level overlap between the synthesized exam-like exercises and the IgakuQA and JMED-LLM test items; substantial overlap would indicate contamination rather than capability.","tokens_in":14584,"feed_emoji":"🩺","tokens_out":7639,"duration_ms":61704,"temperature":0.7,"pith_summary":"This report argues that a carefully built small language model can be competitive with, and on several tasks superior to, multibillion-parameter general models in Japanese clinical and medical text processing. The author constructs a 1.2B-parameter model, NCVC-slm-1, from filtered Japanese Wikipedia and web text, augmented with synthetic medical textbook content, then instruction-tunes it on the eight-task JMED-LLM benchmark. The tuned model records the highest score on six of the eight tasks, including adverse-event classification, symptom detection, sentence similarity, and three named-entity recognition tasks. The reason to care is practical: a model of this size runs on about 2.2 GB of GPU memory, making fully local, privacy-preserving processing of clinical text realistic.","feed_headline":"Small Japanese clinical model tops GPT-4o on 6 of 8 tasks","feed_subtitle":"A 1.2B-parameter model, fine-tuned on Japanese medical benchmarks, runs locally in about 2GB while beating far larger rivals on most tasks.","key_machinery":"The mechanism that carries the result is a domain-tuned tokenization pipeline and a textbooks-style pretraining corpus. Raw text is cleaned, normalized, and then segmented with a Japanese morphological analyzer whose dictionary has been augmented with clinical terms, so medical words survive as whole tokens instead of being fragmented into characters or fallback UTF-8 pieces; a Unigram tokenizer with a 32,768-token vocabulary converts the analyzed text into subword units. The pretraining corpus mixes filtered Wikipedia and Common Crawl text with roughly 0.24B tokens of scraped and synthetic medical textbooks, the synthetic portion being generated by larger language models from disease and drug lists and including exam-like exercises. On top of this, instruction tuning on the JMED-LLM training splits teaches the base model the benchmark's task formats. Each component does specific work: morphological analysis preserves medical vocabulary, the synthetic textbooks supply a domain signal that the 2.6% medical fraction of the corpus would otherwise lack, and instruction tuning converts raw language-modeling competence into usable downstream task performance.","core_discovery":"The paper's central claim is that domain specialization plus instruction tuning can make a 1B-scale model outperform much larger models on a range of Japanese clinical tasks. Concretely, NCVC-slm-1-instruct achieves the top reported score on 6 of the 8 JMED-LLM tasks, namely CRADE, SMDIS, JCSTS, MRNER-disease, MRNER-medicine, and NRNER, beating models such as GPT-4o on those tasks; it remains below GPT-4o on JMMLU-Med and RRTNM, which the author attributes to a higher demand for broad knowledge and inference. The paper also reports that the untuned base model scores near chance on JMED-LLM, which the author reads as evidence that task-specific instruction tuning, rather than scale alone, is what unlocks the small model's performance.","pith_inferences":["The six-task advantage may be partly an artifact of asymmetric training: NCVC-slm-1 was fine-tuned on JMED-LLM training sets, while the comparison scores were quoted from external evaluations of models not similarly fine-tuned; an apples-to-apples run might shrink the margin.","The pretraining corpus contains synthetic exam-like exercises, and the report does not test for overlap with IgakuQA or JMED-LLM; contamination is a testable alternative explanation for part of the result.","A natural extension is to evaluate the model on held-out clinical documents, such as discharge summaries from institutions not represented in the training web text, to see whether the task-format advantage survives outside benchmark prompts.","The architecture choices, including morphological tokenization, grouped query attention, RMSNorm, and SiLU activation, could be ablated independently to identify which one contributes most to the benchmark wins."],"forward_implications":["A 1.2B-parameter model is shown capable of top scores on six JMED-LLM tasks, so local inference on a single GPU is sufficient for several practical clinical text tasks, including named-entity recognition from medical reports and nursing records.","Because fine-tuned small models beat much larger general models on those six tasks, benchmark-specific instruction tuning is the decisive ingredient, not parameter count.","The remaining two tasks, JMMLU-Med and RRTNM, favor GPT-4o, indicating that broad medical knowledge and complex inference are still a scalability limit for 1B-scale models.","Adding synthetic exam data to instruction tuning hurt IgakuQA performance and did not clearly help JMED-LLM, so naive data augmentation is not a reliable path to better clinical question answering.","If deployed, this kind of model could keep sensitive patient text in a local environment rather than sending it to a server."],"supporting_citations":[{"why":"Provides the eight-task JMED-LLM benchmark and the instruction-tuning training splits that transform the base model into the instruct model.","marker":"[63]"},{"why":"Supplies the JMED-LLM scores of the large comparison models that the fine-tuned SLM is claimed to beat on six tasks.","marker":"[67]"},{"why":"Introduces the textbooks approach that motivates filtering pretraining text to high quality and generating synthetic textbooks.","marker":"[26]"},{"why":"Defines the IgakuQA licensing-exam benchmark used to measure the base and instruct models' medical knowledge.","marker":"[65]"},{"why":"Is the cited source for comparison-model scores on IgakuQA, including the top-scoring MedSwallow model.","marker":"[25]"},{"why":"Provides the Chinchilla scaling-law estimate used to choose roughly 20B seen tokens for a 1B-parameter model.","marker":"[55]"},{"why":"Supports the decision to limit repeated epochs and add dropout to avoid degradation from repeating training tokens.","marker":"[56]"}],"fun_headline_variants":["1B Japanese clinical model tops GPT-4o on 6 of 8 tasks","Fine-tuned 1B model beats GPT-4o on 6 Japanese medical tasks","Small Japanese medical model outperforms GPT-4o on 6 benchmarks","1B-scale Japanese clinical SLM beats GPT-4o on 6 of 8 tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the premise that the rival models' quoted scores were produced under comparable evaluation conditions and that fine-tuning on the benchmark's own training data gives the small model no unfair advantage.","fun_headline_variants_meta":{"raw":{"variants":["1B Japanese clinical model tops GPT-4o on 6 of 8 tasks","Fine-tuned 1B model beats GPT-4o on 6 Japanese medical tasks","Small Japanese medical model outperforms GPT-4o on 6 benchmarks","1B-scale Japanese clinical SLM beats GPT-4o on 6 of 8 tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000643,"raw_usage":{"total_tokens":2935,"prompt_tokens":901,"completion_tokens":2034,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":1943}},"tokens_in":517,"tokens_out":2034,"duration_ms":12016,"temperature":1.0,"reasoning_tokens":1943,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:35:41.220890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same eight JMED-LLM tasks under one protocol for all models, using a version of NCVC-slm-1 trained only on the allowed training split and comparison models evaluated in the same harness; if the small model no longer tops six tasks, the central claim is falsified. Separately, search for n-gram or embedding-level overlap between the synthesized exam-like exercises and the IgakuQA and JMED-LLM test items; substantial overlap would indicate contamination rather than capability.","supporting_citations":[{"cited_title":"Jmed-llm: Japanese medical evaluation dataset for large language models","cited_arxiv_id":null,"evidence_quote":"Provides the eight-task JMED-LLM benchmark and the instruction-tuning training splits that transform the base model into the instruct model."},{"cited_title":"https://speakerdeck.com/fta98/ri-ben-yu-yi-liao-llmping- jia-bentimakunogou-zhu-toxing-neng-fen-xi","cited_arxiv_id":null,"evidence_quote":"Supplies the JMED-LLM scores of the large comparison models that the fine-tuned SLM is claimed to beat on six tasks."},{"cited_title":"https://tech.preferred.jp/ja/blog/llama3-preferred-medswallow-70b/","cited_arxiv_id":null,"evidence_quote":"Is the cited source for comparison-model scores on IgakuQA, including the top-scoring MedSwallow model."},{"cited_title":"To repeat or not to repeat: Insights from scaling llm under token-crisis","cited_arxiv_id":null,"evidence_quote":"Supports the decision to limit repeated epochs and add dropout to avoid degradation from repeating training tokens."}],"review_version":1}