{"id":"f2a6724f-1eee-4b46-b045-1d52796a71e8","arxiv_id":"2608.11036","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new Burmese medical speech corpus is released and Whisper fine-tuning (full and LoRA) is benchmarked, with the best model reaching 23.44% WER.","lead":"This paper builds a 28-hour Burmese medical speech corpus from native speaker recordings and fine-tunes Whisper models for clinical dialogue recognition. The best model, a fully fine-tuned Whisper-Medium, reaches 23.44% word error rate on the new test set, and the dataset is planned for public release.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOTA claim rests on a speaker-overlapped test set and ignores the paper's own 19.0% myMediCon baseline; 23.44% is not established as state-of-the-art.","rationale":"Reader's weakest assumption (simulated acoustics) is not the load-bearing point for the headline result. The strongest claim is the SOTA WER; that claim is undermined by the evaluation design (same speakers in train and test, Table 1) and by the paper's own cited 19.0% baseline. The speaker overlap is an internal fact, not a speculation: each speaker row has both training and test utterances. This makes the comparison with zero-shot external baselines unfair and inflates the apparent benefit of fine-tuning. The simulated-acoustics limitation affects only Section 5.2 robustness claims, not the headline 23.44% number. The paper also has secondary arithmetic inconsistencies (e.g., PEFT Small WER 64.80 vs SER+DER+IER=62.67; RTF 'increases from 0.691 to 0.408'), which reduce confidence in reported metrics but are less central than the evaluation-validity issue. The reader's CONDITIONAL verdict is still appropriate: the corpus contribution is real, but the central SOTA claim must be re-evaluated under a speaker-disjoint split and re-benchmarked against myMediCon before acceptance as stated.","tokens_in":11771,"tokens_out":11154,"duration_ms":98301,"concrete_test":"Release or compute a speaker-disjoint evaluation: hold out one or two complete speakers (e.g., sp05 and sp09) from the 52.95-hour training set, fine-tune myMediWhisper-Medium without augmentation under the same hyperparameters, and evaluate on all held-out speaker utterances. If the WER rises materially above 23.44% — especially if it approaches or exceeds the zero-shot whisper-large-v3-myanmar baseline of 32.03% — the headline advantage is attributable to speaker overlap rather than to general domain adaptation. Separately, either report WER on the myMediCon test set (the 19.0% baseline) or remove the unqualified 'state-of-the-art' wording and compare only against Table 2 baselines.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('state-of-the-art WER of 23.44%') requires the evaluation to measure generalization to speakers and to be compared against the best existing result. Table 1 shows every one of the nine speakers appears in both the Train and Test columns (e.g., sp01: 7,364 train / 410 test; sp09: 900 train / 50 test), and the paper never describes a speaker-independent split. Because the comparison systems in Table 2 (chuuhtetnaing/whisper-large-v3-myanmar, MMS-1B, wav2vec2) are zero-shot on this corpus, they have not seen these speaker voices, whereas the FFT and PEFT models have been trained on other utterances from the same speakers. Same-speaker test sets can inflate WER by exploiting speaker-identity cues, so 'outperforming much larger general-domain fine-tuned models' and 'state-of-the-art' are not established for unseen speakers. Independently, Section 1 cites myMediCon's 19.0% WER as the best prior baseline, but the 23.44% result is never benchmarked against it, and the Abstract/Conclusion still use the term 'state-of-the-art.' The result may be a useful domain-adapted Whisper benchmark on the new corpus, but the central claim is over-stated until either a speaker-disjoint evaluation is reported or the SOTA label is qualified against the 19.0% baseline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces myMediWhisper, a new 28-hour Burmese medical speech corpus recorded and validated by native speakers, and uses it to fine-tune Whisper models with full fine-tuning (FFT) and parameter-efficient fine-tuning (PEFT) via rsLoRA. The authors evaluate ASR performance on clean speech and under simulated noise and room acoustics, report that augmentation improves robustness at the cost of clean-speech accuracy, and claim a state-of-the-art WER of 23.44% for the fully fine-tuned myMediWhisper-Medium without augmentation, outperforming larger general-domain fine-tuned models. They also provide a syllable-level error analysis showing voicing/tonal confusions and deletions of weak syllables and grammatical particles.","tokens_in":12083,"tokens_out":5072,"duration_ms":40375,"significance":"If the central claims hold, the paper makes two valuable contributions: it releases a publicly available, native-validated Burmese medical speech corpus, which is rare for this language and domain, and it demonstrates that a mid-size Whisper model can be adapted to a specialized medical domain with limited data. The systematic comparison of FFT and PEFT, the controlled robustness analysis, and the error analysis are useful for practitioners. However, the headline 'state-of-the-art' claim is not currently supported because the evaluation set overlaps in speakers with the training set, and the paper does not benchmark against the 19.0% myMediCon WER it cites as the best prior baseline. The empirical comparison trends are plausible, but the internal inconsistencies in Tables 1 and 2 need to be resolved before the quantitative results can be trusted.","major_comments":[{"comment":"The state-of-the-art claim is not supported by a speaker-independent evaluation. Table 1 lists every speaker in both the Train and Test columns (e.g., sp01: 7,364 train / 410 test; sp09: 900 train / 50 test), and the paper never describes a speaker-disjoint split. Because the fine-tuned models have seen other utterances from the same speakers while the external baselines in Table 2 are zero-shot on this corpus, the comparison conflates domain adaptation with speaker adaptation. The 23.44% WER should also be compared with the 19.0% myMediCon baseline cited in Section 1 before 'state-of-the-art' is used.","section":"§5.1, Table 1"},{"comment":"The error decomposition is internally inconsistent: for several rows SER + DER + IER does not equal the reported WER. For example, PEFT Tiny lists 66.00 + 22.95 + 42.96 = 131.91 but reports WER = 115.68; PEFT Base lists 76.03 + 10.35 + 17.49 = 103.87 but reports 99.57; PEFT Tiny Aug lists 45.90 + 44.76 + 11.79 = 102.45 but reports 92.28. These mismatches indicate a data-processing or reporting error that affects the reliability of all comparisons in this table.","section":"Table 2"},{"comment":"The totals in Table 1 do not match the sum of the row values. The train durations sum to 51.95 hours, not 52.95; the test durations sum to 2.89 hours, not 2.87; and the total durations sum to 54.84 hours, not 55.82. In addition, the text states that the verified corpus contains 28 hours 6 minutes 36 seconds, and the table reports 55.82 total hours 'after adding simulated speech with different room acoustics' (Table 1 caption); the relationship between the original corpus size and the reported training/test durations is not explained.","section":"Table 1"},{"comment":"The claim that models 'generalize reasonably well to unseen room acoustics and noisy conditions' (Section 5.2) is overstated because the robustness test uses the same class of simulated room impulse responses and additive Gaussian noise that were used in the training augmentation. The test conditions are not acoustically independent of the training conditions, and the paper's own Limitations section acknowledges that real-world clinical acoustics may differ. The qualitative conclusion should be framed as robustness to matched simulated variation rather than generalization to unseen acoustic environments.","section":"§5.2, Figures 2-3, §7 Limitations"}],"minor_comments":[{"comment":"The term 'state-of-the-art' is used in the Abstract and Conclusion without comparison to the 19.0% myMediCon baseline or a speaker-independent evaluation; it should be replaced by a qualified claim such as 'best on our evaluation set'.","section":"Abstract, §5.1, §7"},{"comment":"The sentence beginning 'While the open-source fine-tuned whisper-large-v3-myanmar (chuuhtetnaing, 2024) achieves a baseline WER of 32.03% and sil-ai/wav2vec2-bloom-speech-mya (SIL Global - AI, 2022) , our FFT...' is grammatically incomplete; it appears to be missing a verb or comparison for the second model.","section":"§5.1"},{"comment":"The chrF values of 1e-16 for the zero-shot Whisper models are suspiciously near zero and may be a formatting or computation artifact; please report actual chrF scores or explain why they are effectively zero.","section":"Table 2"},{"comment":"The paper reports a verified corpus of 28 hours 6 minutes 36 seconds but later says 52.95 hours of speech were allocated for training; please clarify how the simulated room acoustics expand the data and how the numbers in Figure 1 and Table 1 relate to the original corpus.","section":"§4, Table 1"},{"comment":"The caption does not state which fine-tuning strategy (FFT or PEFT) is used for 'Whisper Medium' and 'Whisper Medium with augmentation'; please specify the configuration to avoid ambiguity.","section":"Figure 3"},{"comment":"The scaling factor in Equation (3) is written as α√r, while the text describes 'a modified scaling factor α/√r'; the notation should be made consistent (e.g., α/√r) and the forward equation should use parentheses to avoid ambiguity.","section":"§3.2, Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern lands: the central SOTA claim is undermined by the speaker-overlapping evaluation and the missing myMediCon comparison. The manuscript also contains several internal numerical inconsistencies in Tables 1 and 2 that must be corrected. The corpus release and the systematic FFT/PEFT comparison are genuine contributions, so I recommend major revision rather than rejection; the authors should be asked to provide a speaker-disjoint evaluation, benchmark against myMediCon, and fix the tables."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the corpus: 28 hours of Burmese medical speech, nine native speakers, syllable-level validation, and a public release on Hugging Face if the link stays live. That is a genuinely reusable resource for a low-resource domain, and the per-speaker stats in Table 1 are useful. The fine-tuning comparison (FFT vs rsLoRA PEFT) and the robustness analysis are standard but competently executed, and the syllable-level error analysis with IPA confusion pairs is a nice touch. Give credit where it is due: the empirical work is honest in its setup, and the paper gives enough detail to reproduce the pipeline.\n\nNow the soft spots, in descending order of seriousness. First, the headline: “state-of-the-art WER of 23.44%” is not established. The paper itself cites myMediCon at 19.0% in Section 1 but never benchmarks against it. You cannot claim SOTA while ignoring a better existing number. Second, the test set is not speaker-independent. Every one of the nine speakers appears in both the Train and Test columns of Table 1, and the paper never mentions a speaker-disjoint split. The external baselines in Table 2 are zero-shot, so they have not seen these voices, while the FFT and PEFT models have trained on other utterances from the same speakers. That gives the fine-tuned models an inherent advantage that is not about domain adaptation alone. The SOTA claim should be withdrawn or heavily qualified until a speaker-independent evaluation is reported.\n\nThird, the tables are sloppy. Several Table 2 rows do not decompose correctly (e.g., PEFT Tiny: 66.00 + 22.95 + 42.96 = 131.91, not 115.68). Table 1 totals do not match row sums: the train utterance column sums to 26,264, not 20,264, and duration sums are off by about an hour. The RTF paragraph contradicts itself: it says PEFT increases RTF but then lists a decrease from 0.691 to 0.408 for Whisper-Medium. These are fixable copy-editing problems, but they are numerous enough to undermine confidence in the reported numbers until corrected.\n\nFinally, the robustness conclusions rest on simulated noise and room impulse responses. The authors acknowledge this as a limitation, and the stress test worried about it. That is a fair limitation, not a fatal flaw — simulated acoustics are common in ASR robustness work. Just do not over-read the results as proven clinical robustness.\n\nIs a serious referee warranted? Yes, because the corpus is valuable and the comparison framework is worth publishing after a major revision. The paper needs a corrected evaluation protocol, a disavowal of the unsupported SOTA claim, and a thorough pass on the tables and text. If those get fixed, this is a solid empirical contribution for the low-resource ASR community.","headline":"Useful new Burmese medical speech resource, but the state-of-the-art claim is not supported by the evidence and the tables need cleaning.","tokens_in":12631,"tokens_out":1902,"would_cite":true,"duration_ms":17886,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A mid-size Whisper fine-tuned on 28 hours of Burmese clinical speech hits 23.44% word error rate, beating far larger general-purpose models.","keywords":["Burmese ASR","medical speech recognition","Whisper fine-tuning","low-resource ASR","data augmentation","simulated room acoustics","parameter-efficient fine-tuning","clinical dialogue"],"falsifier":"Record a separate set of Burmese clinical dialogues in an actual clinic (with real device noise, multi-speaker babble, and microphone characteristics), run myMediWhisper-Medium both with and without augmentation alongside whisper-large-v3-myanmar, and check whether the augmented model beats the unaugmented one and whether myMediWhisper-Medium keeps its WER advantage; if either fails, the paper's robustness-transfer and domain-superiority claims would be contradicted.","tokens_in":11582,"feed_emoji":"🩺","tokens_out":12658,"duration_ms":91064,"temperature":0.7,"pith_summary":"The paper claims that a mid-size Whisper model fully fine-tuned on a 28-hour, native-verified Burmese medical speech corpus reaches a word error rate of 23.44% on clean clinical dialogue, beating much larger general-domain fine-tuned models. The authors argue that domain-specific adaptation is more consequential than raw model scale for specialized medical vocabulary and clinical speaking style, and that the released corpus makes this reproducible. They also show that waveform- and spectrogram-level data augmentation improves robustness to noise and room reverberation at a small cost to clean-speech accuracy, and that parameter-efficient fine-tuning with rank-stabilized LoRA can adapt Whisper-Large-v2 under tight memory constraints. If correct, this provides a concrete recipe for building clinical ASR for low-resource tonal languages.","feed_headline":"Mid-size Whisper beats larger models on Burmese clinic speech","feed_subtitle":"A 28-hour native-verified Burmese medical corpus pushes a mid-size Whisper to 23.44% word error rate.","key_machinery":"The central object is the newly constructed mediTalk-mm-rdy corpus: 14,517 verified Burmese sentences from the Samson PLAB 2 handbook, read by nine native speakers (two male, seven female) and manually checked at syllable level for acoustic fidelity, yielding 28 hours 6 minutes 36 seconds of speech. From this, 52.95 hours are used for training and 2.87 hours for evaluation, and simulated room impulse responses (L-shaped, short/long distance) generated by Pyroomacoustics create reverberant training samples. The fine-tuning pipeline compares full fine-tuning (FFT) with parameter-efficient rank-stabilized LoRA (rsLoRA, $r=128$, $\\alpha=256$, dropout 0.05 applied to the query and value projections) and expands training data 3.75$\\times$ via a 50% sample, three waveform transforms (time shifting, pitch shifting, additive Gaussian noise) and two SpecAugment masks (time and frequency masking). The comparison of WER across Whisper Tiny through Large-v2, with and without augmentation and under controlled SNR and room types, carries the argument that domain-specific fine-tuning on high-quality validated data, not model size alone, drives accuracy.","core_discovery":"On its own terms, the paper establishes that a mid-size Whisper model, fully fine-tuned on a small but carefully validated Burmese medical speech corpus, reaches a state-of-the-art word error rate of 23.44% on a held-out clean test set of clinical dialogues, outperforming general-domain fine-tuned models that are several times larger (whisper-large-v3-myanmar at 32.03% and MMS-1B at 37.90%). This best result is obtained without data augmentation; applying time/pitch shifts, additive Gaussian noise, and SpecAugment masks slightly degrades clean-speech WER (Medium Aug reaches 24.73%) but substantially improves robustness under low signal-to-noise ratios and in simulated reverberant rooms. Under memory constraints, parameter-efficient fine-tuning with rank-stabilized LoRA ($r=128$) adapts Whisper-Large-v2 to 41.57% WER, enabling large-model adaptation on limited hardware at the cost of higher inference latency. Syllable-level error analysis attributes residual errors primarily to voiced/unvoiced plosive confusions and the deletion of weak nominal prefixes and grammatical particles.","pith_inferences":["If the pattern holds, other low-resource tonal languages with a small, natively verified corpus of specialist speech may see similar gains from domain-specific Whisper fine-tuning at mid-size scale, even without large general-domain data.","The clean-versus-robust trade-off suggests a deployment strategy the authors do not explicitly propose: training two copies of the model (one clean, one augmented) and selecting by estimated noise level at runtime.","Because the robustness conclusions rest on simulated acoustics, the imminent plan to collect real clinical recordings will directly test whether the augmentation advantage transfers; until then, the clean-speech WER claim is more firmly evidenced than the robustness claim.","The error analysis hints that explicitly modeling Burmese tonal and glottal features (such as the Asat marker) could further reduce WER, but this design hypothesis is not tested in the paper."],"forward_implications":["The publicly released corpus enables reproducible benchmarking of Burmese medical ASR and downstream clinical speech tasks such as speech translation and medical named entity recognition.","Full fine-tuning of a mid-size Whisper model on a small domain corpus is a practical recipe that outperforms far larger general-domain models, indicating that domain match and data quality can outweigh model scale for specialized dictation.","Parameter-efficient fine-tuning with rsLoRA at $r=128$ unlocks Whisper-Large-v2 adaptation under memory limits, broadening the hardware on which low-resource languages can be served, though inference becomes slower.","The augmentation trade-off yields deployment guidance: use the unaugmented model for controlled environments and the augmented model for noisy or reverberant clinics.","The syllable-level error analysis identifies specific phonetic and orthographic failure modes (voicing contrasts, deletion of weak prefixes/particles) that future work can target with tonal modeling or post-processing."],"supporting_citations":[{"why":"Supplies the Whisper architecture and pretrained models that all fine-tuning experiments build on.","marker":"(Radford et al., 2022)"},{"why":"Prior Burmese medical corpus (myMediCon) and baseline that motivates domain-specific adaptation.","marker":"(Htun et al., 2024)"},{"why":"Source of the verified Burmese translations used as the recording transcripts.","marker":"(Ei San et al., 2022)"},{"why":"Pyroomacoustics library used to simulate room impulse responses for robustness training and evaluation.","marker":"(Scheibler et al., 2018)"},{"why":"LoRA, the low-rank adaptation method that the PEFT experiments extend.","marker":"(Hu et al., 2022)"},{"why":"Rank-stabilized LoRA scaling used in the PEFT configuration.","marker":"(Kalajdzievski, 2023)"},{"why":"SpecAugment provides the spectrogram-level time/frequency masking used in augmentation.","marker":"(Park et al., 2019)"},{"why":"audiomentations supplies the waveform-level transforms (time shift, pitch shift, noise).","marker":"(Jordal et al., 2023)"},{"why":"Whisper-large-v3-myanmar general-domain baseline that myMediWhisper-Medium outperforms.","marker":"(chuuhtetnaing, 2024)"},{"why":"MMS-1B baseline used as external comparison.","marker":"(Pratap et al., 2023)"}],"fun_headline_variants":["Burmese medical speech: 28-hour corpus yields 23.44% WER for Whisper","Native-verified 28-hour corpus pushes mid-size Whisper to SOTA WER","Whisper fine-tuned on Burmese clinic speech hits 23.44% WER","Full fine-tuning beats LoRA for Burmese medical Whisper","Augmentation hurts clean speech but helps noisy Burmese clinic ASR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The robustness conclusions rely on simulated room impulse responses and additive Gaussian noise behaving like real clinical acoustics, so the measured robustness gains are assumed to transfer to actual hospital environments; the authors acknowledge this limitation in Section 7.","fun_headline_variants_meta":{"raw":{"variants":["Burmese medical speech: 28-hour corpus yields 23.44% WER for Whisper","Native-verified 28-hour corpus pushes mid-size Whisper to SOTA WER","Whisper fine-tuned on Burmese clinic speech hits 23.44% WER","Full fine-tuning beats LoRA for Burmese medical Whisper","Augmentation hurts clean speech but helps noisy Burmese clinic ASR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000779,"raw_usage":{"total_tokens":3457,"prompt_tokens":971,"completion_tokens":2486,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2382}},"tokens_in":587,"tokens_out":2486,"duration_ms":16161,"temperature":1.0,"reasoning_tokens":2382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:34:22.945252+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a separate set of Burmese clinical dialogues in an actual clinic (with real device noise, multi-speaker babble, and microphone characteristics), run myMediWhisper-Medium both with and without augmentation alongside whisper-large-v3-myanmar, and check whether the augmented model beats the unaugmented one and whether myMediWhisper-Medium keeps its WER advantage; if either fails, the paper's robustness-transfer and domain-superiority claims would be contradicted.","supporting_citations":[],"review_version":1}