{"id":"46ada474-702e-48f2-986d-edaa7e97dc1e","arxiv_id":"2607.17270","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Medical correctness of small deployable language models drops sharply when questions move from English to Hausa, while a frontier model stays accurate in both.","lead":"Small, locally run medical chatbots gave mostly correct answers in English but often wrong or harmful answers in Hausa. A much larger frontier model stayed accurate in both languages, showing the problem is the class of model, not the language.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Frontier-tier attribution rests on a single API model; without replication across models the 'deployable tier' conclusion may overreach.","rationale":"I read the paper in good faith. The empirical core—five locally deployable models degrading from English to Hausa while one frontier model does not—is internally consistent and carefully reported, including the low harm kappa and the post hoc adjudication rule. The paper is transparent about the study's small scale and its limitations. However, the strongest claim ('It is a property of the deployable tier') goes beyond what the data can establish because it generalizes from a single frontier system to an entire tier. The reader's weakest_assumption identifies exactly this. My concern is not an internal inconsistency; it is an external validity gap that the authors themselves flag. The proposed concrete test—adding more frontier models—would settle whether the frontier behavior is general or idiosyncratic. If multiple frontier models maintain competence, the tier attribution becomes much more credible; if any fails, the conclusion must be weakened to a narrower claim about specific models. This does not change the reader's verdict (CONDITIONAL), as my concern reinforces the conditionality rather than overturning it. The paper's practical recommendation (evaluate at the tier and language of use) remains sound regardless of the frontier generalization, but the specific causal attribution to 'the deployable tier' is what needs the additional evidence.","tokens_in":6989,"tokens_out":3405,"duration_ms":37613,"concrete_test":"Run the identical 12-item English-Hausa evaluation on at least four additional frontier models (e.g., GPT-4o, Claude 3.7 Sonnet, GPT-4.1, Gemini 1.5 Pro, Llama 3.1 405B via API) under the same temperature-zero decoding, and on a broader set of locally deployable models (e.g., Phi-3-mini, Mistral 7B, Llama-3-8B) served through Ollama. If any frontier model exhibits a large Hausa correctness drop (e.g., mean below 1.0), the tier attribution fails. Also attempt to rerun the previously excluded API calls to check whether the frontier's 1.75 mean is sensitive to the exclusion rule. Report per-model trajectories and pooled tier means.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference—that the clinical correctness deficit in Hausa is a property of the deployable tier, not of the language or the clinical material—depends entirely on the premise that a frontier model can answer the same Hausa items competently. That premise is supported by exactly one system (Gemini API). The paper itself admits in the Discussion that 'the frontier tier is represented by one system whose behaviour need not generalise to others.' If this one model happens to have unusually robust Hausa support (e.g., through specific multilingual training or incidental exposure to the exact prompt phrasing), the comparison reduces to 'one large model outperforms five small models,' not a tier-level distinction. The excluded API calls (visible in Figure 7) compound the risk: if service errors occurred disproportionately on difficult Hausa prompts, the frontier's observed correctness would be inflated, and its near-ceiling performance would be an artifact of selective completion. The secondary item-equivalence concern (e.g., the Yoruba 'agbo' item) is real but less decisive, because the frontier model's success shows the items are not fundamentally unanswerable in Hausa. The load-bearing weakness is therefore the representativeness and completeness of the frontier reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates whether clinical safety established in English transfers to Hausa for the class of small, locally deployable language models used in low-resource health settings. Six models were tested: five locally deployable systems (4–9B parameters, two medically fine-tuned) and one frontier API model. Twelve matched English–Hausa items across three conditions (malaria, sickle cell disease, tuberculosis) and four question forms were scored against Nigerian national treatment guidelines by two blind fluent Hausa raters. The central finding is that locally deployable models' mean clinical correctness fell from 1.57 in English to −0.03 in Hausa, while the frontier model moved from 2.00 to 1.75 and produced no harmful responses in either language. The paper concludes that the deficit is a property of the deployable tier, not of the language or the clinical material, and argues that safety assurance must be conducted at the tier and in the language of intended use.","tokens_in":7210,"tokens_out":4264,"duration_ms":41898,"significance":"If the central finding holds, this is an important contribution to multilingual safety evaluation for medical LLMs. The study is anchored to an external standard (Nigerian national treatment guidelines), uses blind dual rating with substantial agreement on the principal endpoint (κ=0.70), and the second rater independently reproduced the tier separation. The authors are notably transparent: they report the poor raw harm κ, describe the post-hoc adjudication rule, flag the single-frontier-model limitation, and release prompts, outputs, and code. The paper identifies a concrete failure mode—correctness drift rather than refusal drift—that is likely to be missed by existing multilingual safety benchmarks. However, the headline attribution to a 'deployable tier' rests on a single frontier model, and the small sample without uncertainty quantification tempers the strength of the conclusion.","major_comments":[{"comment":"The central conclusion—that the deficit is a property of the 'deployable tier'—rests entirely on a contrast between five local models and one frontier API model. The Discussion acknowledges this ('one system whose behaviour need not generalise'), but the Abstract and Conclusion still state the tier-level attribution as a definitive finding. As presented, the evidence supports 'five small models degraded; the one frontier model tested did not', not a general property of all frontier systems. Please either evaluate at least two or three additional frontier models (or show that different frontier models behave similarly on these items), or rephrase the central claim to refer to 'the frontier model in this study' rather than 'the frontier tier'.","section":"Discussion, Table 1, Fig. 7"},{"comment":"A subset of frontier API calls returned service errors and was excluded, but the paper does not report how many calls were excluded, for which items/languages, or whether exclusion was related to prompt difficulty. If service errors occurred disproportionately on longer or more complex Hausa prompts, the frontier's near-ceiling 1.75 mean could be inflated by selective completion. Please provide the excluded-call counts per item and language, and include a sensitivity analysis (e.g., assigning worst-case scores to excluded calls) to show that the tier separation is robust to these missing data.","section":"Results, Fig. 7 and Methods"},{"comment":"The paper reports pooled means without any measure of uncertainty. With only five local models and one frontier model, and with multiple items per model, the difference between −0.03 and 1.75 could be sensitive to item selection or model-specific outliers. Report per-model means (Fig. 3 already does) and provide item-level bootstrap confidence intervals or a mixed-effects model with random effects for item and model. This would allow readers to assess whether the tier contrast is credible beyond this specific item set, rather than relying on the raw mean difference.","section":"Results, Table 1 and Figs. 1–3"}],"minor_comments":[{"comment":"The Dangerous Confidence Rate is defined in Methods but never reported in Results. Either report it for each tier/language or remove the definition to avoid an unused metric.","section":"Methods, Results"},{"comment":"Several statements lack citations: 'Work on Singaporean and Albanian contexts' and the benchmarks 'Med-SafetyBench, CARES' are mentioned without reference numbers. Please add the relevant citations.","section":"Related work"},{"comment":"The reference list contains untracked '[VERIFY: edition and year]' placeholders. These must be replaced with complete bibliographic details before publication.","section":"References [10]–[14]"},{"comment":"The axes 'silent failure' and 'harmful failure' are not formally defined in the text or caption. Define the thresholds used to place models on these axes, since the figure is central to the failure-mode discussion.","section":"Fig. 6"},{"comment":"The abstract states 'All 128 responses were scored', but the total possible responses would be 144 (6 models × 12 items × 2 languages) and some frontier calls were excluded. Clarify how 128 is reached (e.g., excluded calls, timeouts) and state the exact number of excluded calls in the text and figure caption.","section":"Abstract, Fig. 7"}],"recommendation":"major_revision","confidential_remarks":"The paper's core observation is plausible and well presented, and the authors have been unusually honest about limitations. However, the headline claim of a 'deployable-tier' property rests on a single frontier model, which I regard as load-bearing for the conclusion as stated. The additional frontier systems or a suitably weakened claim are needed before publication. I also note that the paper's own limitation paragraph effectively pre-empts this criticism, so the gap is acknowledged; nevertheless, the conclusion as written overreaches."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper, worth a serious referee. The new thing here is real: not another refusal-drift benchmark, but a correctness-drift measurement in Hausa against Nigerian national guidelines, on the small quantised models people actually deploy in low-resource clinics. The result is striking: mean clinical correctness among five local models falls from 1.57 in English to -0.03 in Hausa, crossing into harmful territory, while one frontier model stays near ceiling. The design earns credit: external standard, blind dual raters, and unusually honest reporting—they show the poor harm kappa (0.22), the post hoc adjudication rule, and the admitted Yoruba 'agbo' item.\n\nThe soft spots are the ones the authors name. The tier-level conclusion rests on a single frontier API model, and some of its calls were excluded due to service errors. If another frontier model had failed in Hausa, the comparison reduces to 'one large model beats five small ones.' The paper says this in the Discussion, but the abstract and conclusion state the tier property more confidently than the evidence supports. Second, there are no inferential statistics; the pooled means are directionally consistent across three conditions, and the blind second rater reproduces the tier separation, so I don't think that's fatal—but significance testing would help. Third, item equivalence between English and Hausa is assumed, and the agbo item is an acknowledged breach.\n\nThe citation pattern looks fine, and the released data is a real plus. This is a conditional accept, not a reject. The measurement construct is useful, the finding is actionable if true, and the authors are explicit about what the data can't support. Who is this for: clinical NLP evaluators, AI-safety people concerned with low-resource deployment, and anyone doing multilingual benchmarks. I'd take it seriously in review. Take the tier-level attribution with a grain of salt and push them to add more frontier systems and proper stats, but don't desk-reject.","headline":"Small local medical models drop from competent in English to harmful in Hausa; the tier-level conclusion leans on a single frontier model, but the core finding is real and worth refereeing.","tokens_in":7695,"tokens_out":1302,"would_cite":true,"duration_ms":12733,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Clinical safety measured in English does not transfer to Hausa: locally deployable medical models score as harmful in Hausa while a frontier model stays competent.","keywords":["clinical correctness drift","cross-lingual safety","Hausa","low-resource deployment","local language models","medical safety evaluation","dangerous confidence rate","Nigerian treatment guidelines"],"falsifier":"Run the same twelve matched English-Hausa items under the same rubric with several additional frontier systems, and also re-score the released outputs with clinically qualified raters. If any second frontier model produces a harmful or guideline-violating Hausa answer on the emergency or leading-question items, the claim that the deficit is a property of the deployable tier fails: the result would be a difference among individual models, not between tiers. The released dataset makes this check immediately performable.","tokens_in":6851,"feed_emoji":"🩺","tokens_out":5548,"duration_ms":47361,"temperature":0.7,"pith_summary":"The paper asks whether clinical safety evidence obtained in English and on frontier systems applies to the small, quantised models actually run offline in low-resource settings. It answers no, on the basis of 128 matched English-Hausa responses scored against Nigerian national treatment guidelines. Five locally deployable models fell from a mean clinical correctness of 1.57 in English to -0.03 in Hausa, crossing into the harmful range on average; the single frontier reference model dropped only from 2.00 to 1.75 and never gave a harmful answer. Because every model answered competently in English and the frontier model answered competently in Hausa, the authors attribute the failure to the deployment tier rather than to the language or the clinical task. The result matters because the users most likely to consult these models in Hausa are also the least likely to detect a fluent but wrong answer.","feed_headline":"Local medical models go from safe in English to unsafe in Hausa","feed_subtitle":"Matched questions scored against Nigerian guidelines show the failure belongs to the deployable tier, not the language.","key_machinery":"The instrument is a matched English-Hausa benchmark of twelve items, four each for malaria, sickle cell disease, and tuberculosis, with four question forms per condition: knowledge recall, emergency triage, a leading question inviting a contraindicated action, and a traditional-remedy claim. Correctness is anchored to Nigerian national treatment guidelines, with explicit prohibited statements, so a recommendation of chloroquine or a two-month tuberculosis stop is harmful by rule. The Dangerous Confidence Rate - the proportion of responses that are unhedged, confident, and clinically wrong - is introduced because refusal-based safety metrics cannot distinguish a safe answer from a confidently","core_discovery":"The central claim is that clinical correctness in English does not transfer to Hausa for the deployable tier, and that the locus of failure is the class of model, not the language or the clinical material. In the authors' data, mean correctness among five locally deployable 4-9 billion parameter models fell from 1.57 in English to -0.03 in Hausa, while a frontier system moved from 2.00 to 1.75. Harmful responses among deployable models rose from 5% of English items to 38% of Hausa items under raw flags (25% after adjudication); the frontier model produced none. Drift occurred in every condition - malaria, sickle cell disease, and tuberculosis - and across all five local models, so the result","pith_inferences":["The authors leave implicit that the tier-attribution argument, if sound, generalises to any low-resource language and any safety-critical domain: a safety property verified in one language and tier is not a property of the model but of the evaluation setting.","A direct testable extension is to run the same twelve Hausa items through several additional frontier systems; if any fails the emergency or leading-question items, the 'deployable tier' conclusion would need to be weakened to 'the particular frontier model tested differs from the particular small models tested.'","The observed loss of medical fine-tuning benefit in Hausa suggests a hypothesis the authors could not test: English medical fine-tuning may narrow, rather than broaden, multilingual clinical competence. One could test this by comparing fine-tuned and base versions of the same architecture on the Hausa items.","The adjudication rule that silent failure in an emergency counts as harm could be formalised into an automatic metric, which would make the harm estimate replicable without relying on rater judgment."],"forward_implications":["Safety evaluation must be conducted at the tier of deployment; frontier results do not transfer downward.","Safety evaluation must be conducted in the language of use; English results do not transfer outward.","For clinical applications, safety evaluation must score correctness, not refusal, because a refusal-based test cannot catch fluent wrong answers.","Models below the frontier tier should not be relied on for clinical guidance in languages in which they have not been evaluated in that tier and that language.","Language-routing failures, such as answering Hausa prompts in Swahili, indicate a distinct failure mode with consequences beyond the clinical case."],"fun_headline_variants":["Safety in English doesn't transfer to Hausa for local medical AI","Small medical AI models fail clinical safety in Hausa","Local MedLM: safe in English, unsafe in Hausa","Drift: English-safe medical models turn harmful in Hausa","Model class, not language, drives clinical safety drift"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The attribution of the whole deficit to the deployable tier rests on a single frontier API model being representative of frontier systems, and on the subset of its calls that succeeded being unbiased; if another frontier model answered incorrectly in Hausa, the conclusion would collapse to a one-model versus five-model comparison.","fun_headline_variants_meta":{"raw":{"variants":["Safety in English doesn't transfer to Hausa for local medical AI","Small medical AI models fail clinical safety in Hausa","Local MedLM: safe in English, unsafe in Hausa","Drift: English-safe medical models turn harmful in Hausa","Model class, not language, drives clinical safety drift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3102,"prompt_tokens":870,"completion_tokens":2232,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":2163}},"tokens_in":614,"tokens_out":2232,"duration_ms":13212,"temperature":1.0,"reasoning_tokens":2163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T18:30:20.755701+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same twelve matched English-Hausa items under the same rubric with several additional frontier systems, and also re-score the released outputs with clinically qualified raters. If any second frontier model produces a harmful or guideline-violating Hausa answer on the emergency or leading-question items, the claim that the deficit is a property of the deployable tier fails: the result would be a difference among individual models, not between tiers. The released dataset makes this check immediately performable.","supporting_citations":[],"review_version":1}