{"id":"5e85a7ee-66a5-4b27-8495-660c4614a59c","arxiv_id":"2501.13884","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A finetuned Qwen2-Audio audio LLM achieves high accuracy on several heart murmur features, but the paper's claim of state-of-the-art performance is contradicted by its own tables for grading and murmur classification.","lead":"This paper finetunes an audio language model, Qwen2-Audio, to classify 11 expert-labeled heart murmur features from phonocardiogram recordings. It reports high accuracy on timing, shape, pitch, and quality, but its main claim of beating state-of-the-art systems is contradicted by low grading accuracy and lower murmur classification accuracy in its own results.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own tables contradict the abstract's '8 of 11' claim: on the only two baselines with direct comparison, the model is worse on systolic grading (33.4% vs 96.6%) and murmur W.acc (75.6% vs 83.2%), and no baseline exists for the five diastolic features.","rationale":"The reader's weakest assumption was that Tables 1 and 2 compare fairly with baselines. That is a valid secondary concern, but the decisive issue is internal: even if the baselines are taken exactly as reported, the abstract's count does not match the paper's numbers. The reader's strongest claim already pointed to this contradiction, but the weakest_assumption field focused on cross-paper comparability. My stress-test identifies the internal contradiction as the single load-bearing issue because it is sufficient to reject the central claim regardless of evaluation-protocol differences. The paper also lacks error bars, multiple seeds, and code, but the primary problem is the mismatch between the headline and the results. A simple tally of the tables settles the matter. No other concern is needed.","tokens_in":8590,"tokens_out":3756,"duration_ms":29861,"concrete_test":"Manually construct a table with all 11 features: systolic timing, shape, grading, pitch, quality (Table 1); diastolic timing, shape, grading, pitch, quality (Table 2); and murmur W.acc (Table 2). For each feature, record this work's W.S accuracy and the best baseline accuracy from [12] or [28]. Count the number of features where this work's accuracy is strictly greater than the baseline. If the count is not 8, the abstract's claim is unsupported. Also verify whether [12] or [28] report any diastolic feature results; if not, those features cannot be included in an 'outperforms SOTA' count.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Qwen2-Audio 'outperforms state-of-the-art methods in 8 of the 11 features and performs comparably in the remaining 3.' The only baseline comparisons in the paper are Table 1 (systolic timing, shape, grading, pitch, quality vs Deep CardioSound [12]) and Table 2 (murmur weighted accuracy vs M2D+AST [28] at 83.2%). Table 1 shows this work W.S achieves 100/100/33.4/100/100 versus 96.6/96.3/96.6/96.6/96.5 respectively. Thus the model outperforms on timing, shape, pitch, and quality (4 features) but is far worse on grading (33.4% vs 96.6%). Table 2 shows W.acc of 75.6% vs 83.2%, also worse. The five diastolic features in Table 2 have no baseline numbers; the dash entries indicate no prior method. Therefore, of the 11 features, only 4 can be claimed as outperforming a cited SOTA, and 2 are clearly worse; the remaining 5 have no SOTA comparator. No combination of the reported results yields 8 outperforming plus 3 comparable. The 'remaining 3' is never specified and cannot be identified from the tables. This is an internal inconsistency, independent of any question of evaluation-protocol comparability.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a PCG-analysis system that combines a SSAMBA-based segmentation front-end with a LoRA-finetuned Qwen2-Audio audio-LLM, and evaluates it on 11 expert-labeled murmur features (systolic and diastolic timing, shape, grading, pitch, quality, plus weighted murmur accuracy) from the PhysioNet CirCor DigiScope dataset. Results are reported on a 75/25 patient-level split, with comparisons to Deep CardioSound and M2D+AST and a zero-shot normal/abnormal evaluation on PhysioNet 2016 and Pascal datasets. The abstract and introduction claim that the LLM-based model outperforms state-of-the-art methods in 8 of 11 features and performs comparably in the remaining 3.","tokens_in":8949,"tokens_out":6599,"duration_ms":60385,"significance":"The question is timely: adapting audio-LLMs to phonocardiogram analysis, and especially to non-binary murmur features, is a useful direction and the diastolic features are genuinely under-studied. The paper is empirical only; it provides no code, no release of prompts or hyperparameters, and no machine-checked artifacts. If the reported comparisons were valid, the result would be of interest, but the headline claim is contradicted by the paper's own tables and the evaluation lacks basic statistical safeguards, so the positive significance is not established in the current form.","major_comments":[{"comment":"The claim that the model 'outperforms state-of-the-art methods in 8 of the 11 features and performs comparably in the remaining 3' is contradicted by the reported numbers. Table 1 shows that for systolic features this work W.S achieves 100/100/33.4/100/100 versus Deep CardioSound's 96.6/96.3/96.6/96.6/96.5, so only timing, shape, pitch, and quality outperform, while grading is 33.4% versus 96.6%. Table 2 shows murmur W.acc of 75.6% versus 83.2% for M2D+AST, also worse. The five diastolic features have no baseline entries. No subset of the tables yields 8 outperforming and 3 comparable features; the 'remaining 3' is never identified. This internal inconsistency must be corrected or the claim removed.","section":"Abstract; §1; Tables 1–2"},{"comment":"The evaluation reports only point estimates of accuracy with no error bars, confidence intervals, statistical tests, or majority-class baselines. Given the small test subsets for long-tail diastolic features and the 33.4% systolic grading result, it is impossible to tell whether differences such as 100% versus 99.7% are meaningful or whether grading performance is at or below chance. The statement in §4 that segmentation 'does improve diastolic grading performance above chance level' is unsupported because no chance baseline is provided.","section":"§4; Tables 1–2"},{"comment":"The comparisons with Deep CardioSound [12] and M2D+AST [28] are not established as protocol-comparable. The paper uses a custom 75/25 patient-level split of CirCor, whereas the cited W.acc definition in [36] and the M2D+AST result in [28] are tied to the PhysioNet Challenge 2022 evaluation setup, and Deep CardioSound [12] was evaluated under its own data split. Without re-running the baselines under identical splits, feature label sets, and accuracy definitions, the claimed outperformance on the four systolic features is unsupported.","section":"§3.3; Tables 1–2"},{"comment":"The absence of any previous results for the five diastolic features does not establish that this model succeeds on them; it only means no comparison was made. The text says these are 'long-tail' features and that previous methods 'failed to classify' them, but no class-frequency statistics or prior negative results are cited. The paper should either provide a baseline trained on the same features or explicitly frame these numbers as first unreplicated measurements rather than as evidence of superiority.","section":"§4; Table 2"}],"minor_comments":[{"comment":"The text says the data are resampled to '16,000 kHz'; this should be 16 kHz.","section":"§3.3"},{"comment":"LoRA is low-rank adaptation, not 'full low-rank approximation'; also 'matrics' should be 'matrices'.","section":"§3.1.2"},{"comment":"There is a typo: 'repeatative' should be 'repetitive'.","section":"§1"},{"comment":"The heading says 'weighed murmur classification'; this should be 'weighted murmur classification'.","section":"Table 2 heading"},{"comment":"The description of SSAMBA finetuning says the model is trained to segment 'occurrences and silences', but it is unclear what labels are used and whether they come from the CirCor heartbeat annotations or are produced by another method; this should be specified for reproducibility.","section":"§3.1.1"},{"comment":"The paper does not report LoRA rank and alpha, learning rate, number of epochs, prompt templates, or the three question phrasing variants, which makes the experiment difficult to reproduce.","section":"§3.1.2 and §3.3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the abstract's central quantitative claim is refuted by the paper's own tables, and the evaluation has no uncertainty quantification. I recommend major revision rather than outright rejection because the underlying exploration of audio-LLMs for multi-feature PCG classification is a reasonable topic and the negative results on grading and weighted accuracy could be reported honestly. However, the authors must substantially reframe the claims, add proper baselines or clearly label the diastolic results as unreplicated, and provide statistical support before the paper can be considered further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is the first to finetune an audio LLM (Qwen2-Audio) for multi-feature heart murmur classification, and it does include diastolic features that prior work skipped. That part is genuinely new. But the headline claim is not supported by the paper's own tables: they say they outperform SOTA in 8 of 11 features and are comparable in 3, yet Table 1 shows they beat Deep CardioSound on only four systolic features (timing, shape, pitch, quality) and are far worse on grading (33.4% vs 96.6%). Table 2 shows they are worse on murmur weighted accuracy (75.6% vs 83.2% for M2D+AST), and the five diastolic features have no baseline numbers at all. So no counting gets you 8+3. The \"remaining 3\" is never identified. This is an internal inconsistency, not just a matter of protocol interpretation.\n\nWhat's good: the application is new, the segmentation front-end seems to help (diastolic grading 39.5->42.7, W.acc 63.7->75.6), and the zero-shot results on external datasets show at least partial robustness. They also honestly admit the grading problem is likely due to the frozen text encoder and Roman numeral labels.\n\nSoft spots: no error bars, no majority baseline, no statistical tests, and the comparison to Deep CardioSound may use different splits or label definitions. The diastolic features have no prior baseline, so \"successful classification\" is just absolute accuracy on a test set with no calibration. The 100% numbers on timing/shape/pitch/quality are suspicious and might reflect test set imbalance or leakage—they do randomize question templates, which is good, but the accuracy is suspiciously perfect.\n\nWho this is for: researchers working at the intersection of audio foundation models and medical acoustics, plus the PCG benchmarking community. They would get a useful case study and a cautionary tale about overclaiming.\n\nRecommendation: this deserves a serious referee, but not in its current form. The core idea is worth reviewing; the claims need to be rewritten to match the evidence, and the evaluation needs uncertainty quantification and a stronger baseline setup. A careful referee could help turn this into a useful benchmark paper. I would not cite it as-is.","headline":"First audio-LLM PCG feature study, but the abstract's '8 of 11 outperformance' is contradicted by the paper's own tables.","tokens_in":9397,"tokens_out":2304,"would_cite":false,"duration_ms":20346,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single finetuned audio LLM, Qwen2-Audio, can classify all 11 clinically relevant heart murmur features and outperform prior specialized models on most of them.","keywords":["audio large language model","heart murmur","phonocardiogram","murmur feature classification","Qwen2-Audio","SSAMBA segmentation","zero-shot classification","low-rank adaptation"],"falsifier":"Re-run Deep CardioSound and M2D+AST on the exact test split of the CirCor DigiScope data used here; if either baseline matches or exceeds the LLM's accuracy on the eight timing, shape, pitch, and quality features, the claimed 8-of-11 outperformance is unsupported.","tokens_in":8431,"feed_emoji":"🩺","tokens_out":10133,"duration_ms":77735,"temperature":0.7,"pith_summary":"This paper tries to establish that a general-purpose audio large language model, Qwen2-Audio, can be finetuned on heart sound recordings to classify a comprehensive set of 11 expert-labeled murmur features—timing, shape, pitch, quality, and grading in both systolic and diastolic phases, plus weighted murmur status—using a single model rather than a battery of specialist classifiers. The authors argue that this approach outperforms existing deep learning baselines on 8 of the 11 features, performs comparably on the remaining 3, and, with the help of a SSAMBA-based segmentation front-end, also generalizes to unseen heart sound datasets in zero-shot normal/abnormal classification. The clinical motivation is that these murmur features are what physicians use to narrow a differential diagnosis, while most prior automated systems only output a healthy-versus-unhealthy label. The paper further claims that the model is the first to classify long-tail diastolic murmur features with limited training data.","feed_headline":"Audio LLM beats prior models on 8 of 11 murmur features","feed_subtitle":"One Qwen2-Audio model, adapted with LoRA and a segmentation front-end, labels systolic and diastolic murmur traits.","key_machinery":"The load-bearing machinery is Qwen2-Audio, an audio LLM whose Whisper-based audio encoder turns a phonocardiogram into representations that a Qwen-7B language model conditions on to produce text answers in a multiple-choice format; it is adapted to the medical domain by low-rank adaptation (LoRA) of both the encoder and the LLM weights. A second component, SSAMBA, is a Mamba state-space audio representation model finetuned with a linear head to segment each recording into heartbeat and non-heartbeat intervals, and feeding these segments to the LLM is what gives the system robustness on unseen datasets. The multiple-choice question format, with randomly varied phrasing of each question, is the mechanism that converts the LLM's generative language modeling into a classifier over the fixed label sets of each of the 11 tasks.","core_discovery":"The paper's central claim is that an audio LLM finetuned on heart sounds can jointly classify the full set of clinically used murmur features, going beyond the binary healthy-versus-unhealthy task that dominates prior work. On its test split of the CirCor DigiScope data, the authors report that the model reaches 100% accuracy on timing, shape, pitch, and quality for both systolic and diastolic murmurs when paired with the SSAMBA segmentation front-end, and that it is the first system to classify long-tail diastolic features at all. The same front-end enables zero-shot normal-versus-abnormal classification on two external heart sound datasets. The authors state the overall result as outperforming state-of-the-art methods on 8 of the 11 features and performing comparably on the remaining 3, while noting that grading features remain difficult for the model.","pith_inferences":["A direct head-to-head rerun of the two baseline systems on the exact test split used here would settle whether the 8-of-11 claim is robust, because the tables compare against numbers reported under different evaluation protocols.","The poor grading results and the paper's own explanation point to a testable fix: unfreezing the text encoder or replacing Roman-numeral labels with numeric ones may recover grading accuracy, and if so the same model could plausibly reach all 11 features.","The segmentation front-end's success on heart sounds suggests a transferable recipe for other periodic biomedical sounds, such as respiratory or bowel sounds, where segmenting events before an LLM could improve few-shot classification.","If audio LLMs can reliably describe murmur features, the next natural step—implicit in the paper but not tested—is to have the same model generate a free-text clinical impression from the murmur description, making it an assistant rather than a labeler."],"forward_implications":["If the central claim holds, a single audio LLM can replace a collection of task-specific deep networks for heart murmur phenotyping, simplifying automated auscultation analysis.","The demonstrated classification of long-tail diastolic features, which previous systems could not label, would give clinicians a machine-readable description of diastolic murmurs that automated tools currently lack.","The segmentation front-end's improvement in zero-shot transfer suggests that preprocessing periodic biomedical audio into meaningful segments is a broadly useful step for audio LLMs, not just for heart sounds.","The joint modeling of 11 features implies that the LLM captures shared acoustic structure across murmur traits, allowing one system to produce a complete murmur description rather than separate binary calls."],"supporting_citations":[{"why":"Qwen2-Audio is the base audio LLM that the paper finetunes; it supplies the audio encoder and the language model.","marker":"[5]"},{"why":"The CirCor DigiScope dataset provides the 11 expert-labeled murmur features and the weighted-accuracy metric used for evaluation.","marker":"[36]"},{"why":"Deep CardioSound supplies the systolic feature accuracies that the paper compares against in Table 1.","marker":"[12]"},{"why":"M2D + AST supplies the 83.2% murmur weighted accuracy that the paper compares against in Table 2.","marker":"[28]"},{"why":"SSAMBA provides the pretrained Mamba audio representation model that becomes the PCG segmentation front-end.","marker":"[37]"},{"why":"LoRA is the low-rank adaptation method used to finetune Qwen2-Audio without full fine-tuning.","marker":"[13]"}],"fun_headline_variants":["Audio LLM tops 8 of 11 heart murmur features","Audio LLM first to classify rare murmur traits","Qwen2-Audio labels all murmur features, bests prior art","Finetuned audio LLM beats prior models on 8/11 features","LLM handles all murmur features, not just healthy vs unhealthy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the baseline accuracies from Deep CardioSound and M2D+AST were measured under the same feature definitions, accuracy metric, patient splits, and label sets as this paper's evaluation, so the numbers in Tables 1 and 2 are directly comparable.","fun_headline_variants_meta":{"raw":{"variants":["Audio LLM tops 8 of 11 heart murmur features","Audio LLM first to classify rare murmur traits","Qwen2-Audio labels all murmur features, bests prior art","Finetuned audio LLM beats prior models on 8/11 features","LLM handles all murmur features, not just healthy vs unhealthy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000952,"raw_usage":{"total_tokens":4080,"prompt_tokens":984,"completion_tokens":3096,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":3008}},"tokens_in":600,"tokens_out":3096,"duration_ms":19309,"temperature":1.0,"reasoning_tokens":3008,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:29:32.269102+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run Deep CardioSound and M2D+AST on the exact test split of the CirCor DigiScope data used here; if either baseline matches or exceeds the LLM's accuracy on the eight timing, shape, pitch, and quality features, the claimed 8-of-11 outperformance is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The CirCor DigiScope dataset provides the 11 expert-labeled murmur features and the weighted-accuracy metric used for evaluation."},{"cited_title":"Deep CardioSound-An Ensembled Deep Learning Model for Heart Sound MultiLabelling","cited_arxiv_id":"2204.07420","evidence_quote":"Deep CardioSound supplies the systolic feature accuracies that the paper compares against in Table 1."},{"cited_title":"Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection","cited_arxiv_id":"2404.17107","evidence_quote":"M2D + AST supplies the 83.2% murmur weighted accuracy that the paper compares against in Table 2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"LoRA is the low-rank adaptation method used to finetune Qwen2-Audio without full fine-tuning."}],"review_version":1}