Pith. sign in

REVIEW 4 major objections 6 minor 36 references

No vision or omni model yet reads fine acoustic detail from medical spectrograms; the best reaches only 51% on CaReCoS.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 03:02 UTC pith:SF22N6DW

load-bearing objection Useful first spectrogram-image QA benchmark for auscultation; the low explicit accuracies are the real signal, while the high inferred scores and overall ceiling rest on unvalidated Gemini labels and an LLM judge. the 4 major comments →

arxiv 2607.03356 v1 pith:SF22N6DW submitted 2026-07-03 eess.AS

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds

classification eess.AS
keywords spectrogrammedical audiocardiac auscultationrespiratory soundscough detectionvision-language modelsmultimodal benchmarkmel spectrogram
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Heart, lung, and cough sounds carry diagnostic information that clinicians often inspect as spectrograms, yet no prior benchmark tests whether multimodal models can reason over those images. The authors build CaReCoS by converting seven public clinical audio datasets into mel spectrograms and pairing each with clinically grounded questions: explicit items that retrieve labeled attributes and inferred items that require multi-step medical reasoning. Nine state-of-the-art vision and omni models receive only the spectrogram image plus the question text. All fail to combine visual pattern recognition with medical knowledge; GPT-5.1 tops the table at 48.8% weighted accuracy and 51.2% on one subset. Medically fine-tuned models lag further behind, exposing a domain mismatch between anatomical image training and time-frequency representations. The result shows that current systems cannot yet extract the fine acoustic features clinicians rely on, so targeted training on medical sound visualizations is required.

Core claim

State-of-the-art vision-language and omni models cannot reliably extract fine-grained acoustic features from mel spectrograms of cardiac, respiratory, and cough sounds and map them to clinical conclusions. On CaReCoS the strongest model, GPT-5.1, reaches only 48.8% weighted accuracy overall and a peak of 51.2% on ICBHI; medically specialized models perform worse, confirming that anatomical vision training does not transfer to time-frequency medical sound patterns.

What carries the argument

CaReCoS, a two-tiered spectrogram QA benchmark that pairs mel spectrograms from seven clinical audio datasets with explicit questions anchored to metadata and inferred questions demanding multi-step clinical reasoning, evaluated under an LLM-as-a-judge protocol.

Load-bearing premise

The paper treats Gemini-generated inferred answers and the LLM-as-a-judge scores as reliable ground truth even though no clinicians skilled in both auscultation and spectrogram reading reviewed them at scale.

What would settle it

Have a panel of clinicians re-annotate a stratified sample of the inferred questions and re-score the model answers; if the top model then exceeds roughly 70% accuracy under the clinician labels, the claim that models cannot combine spectrogram reading with medical knowledge is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Multimodal medical systems will need explicit training on spectrogram visualizations rather than relying on transfer from radiology or pathology images.
  • Medical domain fine-tuning confined to anatomical imagery will not close the gap for auscultation tasks.
  • Spectrogram-based evaluation can expose failures of visual feature extraction that raw-audio or text-only medical knowledge cannot reveal.
  • Pediatric cardiac spectrograms remain especially hard, indicating a need for age-specific acoustic training data.
  • Strict factual matching is required for clinical QA scoring because surface similarity rewards fluent but incorrect answers.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Fine-tuning the same models on paired spectrogram–clinical-text data would likely raise explicit-question accuracy faster than inferred-question accuracy, isolating the perceptual bottleneck.
  • The jump some omni models show when given raw audio instead of spectrograms suggests current vision encoders lack the time-frequency inductive biases clinicians use.
  • Applying the same explicit-versus-inferred design to other scientific visualizations such as ECGs or spirometry traces could reveal a broader visualization-literacy deficit in foundation models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces CaReCoS, a spectrogram-image QA benchmark for medical auscultation sounds drawn from seven public cardiac, respiratory, and cough datasets. Each recording is converted to a mel spectrogram and paired with two metadata-anchored (explicit) and two multi-step clinical (inferred) questions; inferred answers are generated by Gemini 3 Flash from metadata and waveform. Nine vision and omni models are evaluated under an LLM-as-a-judge protocol. The central claim is that no model reliably combines spectrogram pattern recognition with medical knowledge: GPT-5.1 reaches at most 51.2% on ICBHI and 48.8% weighted average overall, with medically fine-tuned VLMs lagging general-purpose models, and with a large inferred–explicit accuracy gap (e.g., 84.6% vs 13.0% for GPT-5.1).

Significance. If the evaluation is reliable, CaReCoS would be a useful and timely contribution: no existing audio or VLM benchmark targets medical sound reasoning via spectrograms, the clinical modality clinicians actually inspect. The multi-dataset coverage, explicit/inferred split, audio-vs-spectrogram ablation, and qualitative analysis of fluent-but-wrong answers are concrete strengths that would help the community measure and improve medical time–frequency understanding. The negative result—that current VLMs and medical VLMs fail at fine-grained spectrogram features—would motivate targeted pretraining and would be of interest to both medical AI and multimodal evaluation venues. Those conclusions, however, currently rest on unvalidated inferred ground truth and an unvalidated LLM judge, so the headline ceiling numbers cannot yet be treated as a clean measure of clinical spectrogram competence.

major comments (4)
  1. Dataset Design and Experiments: Inferred QA ground truth is produced by Gemini 3 Flash from metadata+waveform, and the paper states that clinician validation (auscultation + spectrogram skill) was not performed at scale. Table 2 shows that the reported overall accuracies are driven almost entirely by inferred items (GPT-5.1 84.6% inferred vs 13.0% explicit). Without expert adjudication of a substantial sample of inferred answers, the claim that models fail at “clinical reasoning over spectrograms” is not load-bearing; the high inferred scores may largely reflect language priors and shared clinical vocabulary rather than spectrogram reading. A clinician-reviewed subset (or a clear reframing that treats only explicit items as primary) is needed.
  2. Experiments (judgment stage) and Discussion: Free-form answers are scored by an LLM-as-a-judge with relaxed semantic criteria for inferred questions. The paper correctly notes that embedding similarity fails on fluent but clinically wrong answers, yet it reports no human–judge agreement, no adjudication of judge decisions, and no sensitivity analysis of the judge prompt. Because the same model family used for generation also appears among evaluated systems, shared generation/judgment artifacts could inflate or distort the ceiling. Inter-rater agreement against clinicians (or at least against independent expert raters) on a stratified sample is required before the 48.8%/51.2% numbers can support the abstract’s claim.
  3. Table 2 and Discussion: The explicit results (near-floor for all models, ~7–15%) are the strongest evidence that models cannot extract fine-grained acoustic attributes (murmur duration, grade, quality, location) from mel spectrograms. The abstract and conclusions, however, lead with the mixed overall accuracy and the “clinical reasoning” framing. The manuscript should either (i) make explicit accuracy the primary reported metric and treat inferred accuracy as secondary/exploratory, or (ii) validate the inferred track as above. As written, the headline maximum of 51.2% overstates spectrogram competence relative to the metadata-anchored evidence.
  4. Experiments ablation and Table 1: The audio-vs-spectrogram ablation is valuable but incomplete. Gemini 2.5 Flash’s large gain with raw audio is attributed mainly to a 42% spectrogram refusal rate; refusal handling, prompt format, and whether vision models were allowed to abstain are not fully specified. Also, ZCH has only 80 items and is the hardest set for every model—weighted averages are dominated by CirCor/Coswara. Report refusal rates for all models, confidence intervals or bootstrap estimates, and per-dataset n clearly enough that the weighted average cannot be misread as uniform difficulty.
minor comments (6)
  1. Table 1 caption: “benchamrking resuts” → “benchmarking results”; also “LLaV A” spacing is inconsistent throughout (LLaVA).
  2. Figure 2 caption is incomplete (“relative accuracy of GPT 5.1”) and does not explain the dual axes or what the bars represent.
  3. §3: Clarify total number of unique recordings vs QA pairs, and whether any de-duplication was applied across the seven source datasets.
  4. §4: Mel-spectrogram parameters (n_fft=1024, hop=512, n_mels=256) are fixed with no sensitivity check; a short note on whether results change under common clinical display settings would help reproducibility.
  5. References: Rocha et al. 2019a/2019b appear duplicated; clean the bibliography.
  6. Naming: Generation uses “Gemini 3 Flash” while evaluation uses “Gemini 2.5 Flash/Pro”—confirm model identifiers for reproducibility, as naming conventions change quickly.

Circularity Check

0 steps flagged

No circular derivation: empirical benchmark accuracies are measured against disclosed (partly synthetic) labels, not forced by construction or self-citation.

full rationale

CaReCoS is a benchmark construction and zero-shot evaluation paper, not a first-principles derivation or predictive theory. Explicit QA ground truth is taken directly from the seven source datasets' metadata (ICBHI, CirCor, etc.); inferred QA pairs are generated by Gemini 3 Flash from the same metadata plus waveform, with the paper explicitly stating that clinician validation was not performed at scale. Model accuracies (max 51.2 % on ICBHI, 48.8 % weighted) and the inferred/explicit split are then obtained by feeding mel-spectrograms to nine external VLMs/omni models and scoring free-form answers with an LLM-as-a-judge under stated criteria. None of these steps reduces a claimed prediction or uniqueness result to its own inputs by definition, fit, or load-bearing self-citation. The authors do not invoke prior uniqueness theorems of their own, do not rename known empirical patterns as new theory, and do not present the measured accuracies as theoretically forced. Methodological concerns about synthetic labels and LLM judging affect validity, not circularity of a derivation chain. Score 0 is therefore the correct, non-manufactured outcome.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 1 invented entities

The central negative claim rests on treating LLM-generated inferred answers and LLM judges as proxies for clinical correctness, on mel spectrograms as a fair visual interface, and on the chosen seven public datasets as representative. Spectrogram hyperparameters and the two-tier question design are free design choices that shape measured accuracy. No new physical entities are postulated; the invented object is the benchmark itself.

free parameters (2)
  • mel spectrogram n_fft / hop_length / n_mels
    Fixed at 1024 / 512 / 256 without ablation; these choices control time-frequency resolution and thus what fine-grained features models can see.
  • two explicit + two inferred QA pairs per recording
    Hand-chosen generation quota that defines the benchmark mix and the weighted averages reported.
axioms (4)
  • domain assumption Mel spectrograms of auscultation audio encode the clinically relevant acoustic features needed to answer both explicit and inferred questions.
    Assumed throughout the evaluation protocol (Introduction, Experiments); clinical practice uses spectrograms but not necessarily this exact mel rendering.
  • ad hoc to paper Gemini 3 Flash can generate valid inferred clinical ground-truth answers from metadata and waveform without clinician verification.
    Stated in Dataset Design; paper acknowledges clinician validation is hard to source at scale but still uses these labels as ground truth.
  • ad hoc to paper LLM-as-a-judge with category-specific criteria is a faithful substitute for expert scoring of free-form medical answers.
    Experiments section; qualitative critique of BERT similarity is given, but no clinician agreement study for the judge itself.
  • domain assumption Zero-shot prompting with short free-form answers is a fair test of whether models can combine visual spectrogram reading with medical knowledge.
    Evaluation design in Experiments; alternative protocols (multiple choice, fine-tuning, native audio) are only partially ablated.
invented entities (1)
  • CaReCoS benchmark (spectrogram-image + explicit/inferred clinical QA pairs) no independent evidence
    purpose: Provide a structured multimodal evaluation resource for medical sound reasoning via spectrograms.
    Constructed artifact of the paper; independent evidence would be public release plus external clinician validation, which the text does not confirm.

pith-pipeline@v1.1.0-grok45 · 12949 in / 3236 out tokens · 28694 ms · 2026-07-12T03:02:50.218191+00:00 · methodology

0 comments
read the original abstract

Medical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reasoning over their spectrogram representations. We address both gaps with CaReCoS, a benchmark pairing clinically grounded questions with mel-spectrogram images derived from seven medical audio datasets. Evaluating 9 state-of-the-art vision and omni models, we find that all struggle with fine-grained acoustic features encoded in spectrograms: no model reliably combines visual pattern recognition with medical knowledge, achieving a maximum accuracy of 51.2%, underscoring the need for training on medical sound visualizations.

Figures

Figures reproduced from arXiv: 2607.03356 by Abhishek Mukherji, Akhil Pothanapalli, Asif Shaik, Harshit Rajgarhia, Prasanna Desikan, Rachuri Lokesh, Shuubham Ojha.

Figure 1
Figure 1. Figure 1: A flow diagram of the proposed pipeline for benchmark generation. as a primary input, despite spectrograms being the standard visualization in clinical decision-support software (Fraiwan et al., 2021; Kandasamy et al., 2011). 2.2. General Image Reasoning: Models and Benchmarks Vision-language models (VLMs) have scaled rapidly through instruction-tuning on web-scale image-text data, from LLaVA (Liu et al., … view at source ↗
Figure 2
Figure 2. Figure 2: Number of total QA pairs in CaReCoS and relative accuracy of GPT 5.1 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 9 linked inside Pith

  1. [1]

    2024 , url =

    Tang, Changli and Yu, Wenyi and Sun, Guangzhi and Chen, Xianzhao and Tan, Tian and Li, Wei and Lu, Lu and Ma, Zejun and Zhang, Chao , booktitle =. 2024 , url =

  2. [2]

    arXiv preprint arXiv:2311.07919 , year =

    Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models , author =. arXiv preprint arXiv:2311.07919 , year =

  3. [3]

    arXiv preprint arXiv:2407.10759 , year =

    Qwen2-Audio Technical Report , author =. arXiv preprint arXiv:2407.10759 , year =

  4. [4]

    arXiv preprint arXiv:2503.20215 , year =

    Qwen2.5-Omni Technical Report , author =. arXiv preprint arXiv:2503.20215 , year =

  5. [5]

    , booktitle =

    Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F. , booktitle =. 2025 , url =

  6. [6]

    and Tyagi, Utkarsh and Kumar, Sonal and Seth, Ashish and Selvakumar, Ramaneswaran and Nieto, Oriol and Duraiswami, Ramani and Ghosh, Sreyan and Manocha, Dinesh , booktitle =

    Sakshi, S. and Tyagi, Utkarsh and Kumar, Sonal and Seth, Ashish and Selvakumar, Ramaneswaran and Nieto, Oriol and Duraiswami, Ramani and Ghosh, Sreyan and Manocha, Dinesh , booktitle =. 2025 , url =

  7. [7]

    Yang, Qian and Xu, Jin and Liu, Wenrui and Chu, Yunfei and Jiang, Ziyue and Zhou, Xiaohuan and Leng, Yichong and Lv, Yuanjun and Zhao, Zhou and Zhou, Chang and others , journal =

  8. [8]

    arXiv preprint arXiv:2303.08774 , year =

  9. [9]

    arXiv preprint arXiv:2410.21276 , year =

  10. [10]

    Advances in Neural Information Processing Systems (NeurIPS) , volume =

    Visual Instruction Tuning , author =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  11. [11]

    arXiv preprint arXiv:2502.13923 , year =

    Qwen2.5-. arXiv preprint arXiv:2502.13923 , year =

  12. [12]

    arXiv preprint arXiv:2507.06261 , year =

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author =. arXiv preprint arXiv:2507.06261 , year =

  13. [13]

    Yue, Xiang and Ni, Yuansheng and Zhang, Kai and Zheng, Tianyu and Liu, Ruoqi and Zhang, Ge and Stevens, Samuel and Jiang, Dongfu and Ren, Weiming and Sun, Yuxuan and others , booktitle =

  14. [14]

    Yue, Xiang and others , journal =

  15. [15]

    Liu, Yuan and Duan, Haodong and Zhang, Yuanhan and Li, Bo and Zhang, Songyang and Zhao, Wangbo and Yuan, Yike and Wang, Jiaqi and He, Conghui and Liu, Ziwei and Chen, Kai and Lin, Dahua , booktitle =

  16. [16]

    Fu, Chaoyou and Chen, Peixian and Shen, Yunhang and Qin, Yulei and Zhang, Mengdan and Lin, Xu and Yang, Jinrui and Zheng, Xiawu and Li, Ke and Sun, Xing and others , journal =

  17. [17]

    Li, Baiqi and others , booktitle =

  18. [18]

    Physiological Measurement , volume =

    An Open Access Database for the Evaluation of Respiratory Sound Classification Algorithms , author =. Physiological Measurement , volume =. 2019 , publisher =

  19. [19]

    Oliveira, Jorge and Renna, Francesco and Costa, Paulo Dias and Nogueira, Marcelo and Oliveira, Cristina and Ferreira, Carlos and Jorge, Al. The. IEEE Journal of Biomedical and Health Informatics , volume =. 2022 , publisher =

  20. [20]

    2024 , publisher =

    Zhang, Yuhang and Zhang, Qing and Zhang, Jing and Yuan, Jiajun and Huang, Huajie and Zhang, Baoqin and Lv, Gaomei and Lin, Shuzhu and Wang, Na and Liu, Xin and others , journal =. 2024 , publisher =

  21. [21]

    Data in Brief , volume =

    A Dataset of Lung Sounds Recorded from the Chest Wall Using an Electronic Stethoscope , author =. Data in Brief , volume =. 2021 , publisher =

  22. [22]

    International Journal of Telemedicine and Applications , volume =

    Monitoring and Analysis of Lung Sounds Remotely , author =. International Journal of Telemedicine and Applications , volume =. 2011 , publisher =

  23. [23]

    2022 , publisher =

    Zhang, Qing and Zhang, Jing and Yuan, Jiajun and Huang, Huajie and Zhang, Yuhang and Zhang, Baoqin and Lv, Gaomei and Lin, Shuzhu and Wang, Na and Liu, Xin and others , journal =. 2022 , publisher =

  24. [24]

    and Ghosh, Prasanta Kumar and Ganapathy, Sriram , booktitle =

    Sharma, Neeraj Kumar and Krishnan, Prashant and Kumar, Rohit and Ramoji, Shreyas and Chetupalli, Srikanth Raj and Nirmala, R. and Ghosh, Prasanta Kumar and Ganapathy, Sriram , booktitle =. 2020 , doi =

  25. [25]

    A Multimodal Dataset for Automatic

    Orlandic, Lara and Thevenot, J. A Multimodal Dataset for Automatic. 2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) , pages =. 2023 , publisher =

  26. [26]

    and Maimbolwa, Minyoi M

    Baur, Sebastien and Nabulsi, Zaid and Weng, Wei-Hung and Garrison, Jake and Blankemeier, Louis and Fishman, Sam and Chen, Christina and Kakarmath, Sujay S. and Maimbolwa, Minyoi M. and Sanjase, Nsala and others , journal =

  27. [27]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  28. [28]

    Zhang, Yuwei and Xia, Tong and Saeed, Aaqib and Mascolo, Cecilia , journal =

  29. [29]

    Exploring Fine-Tuned Audio-

    Anonymous , journal =. Exploring Fine-Tuned Audio-

  30. [30]

    arXiv preprint arXiv:2504.07491 , year =

    Kimi-. arXiv preprint arXiv:2504.07491 , year =

  31. [31]

    Li, Bo and Zhang, Yuanhan and Guo, Dong and Zhang, Renrui and Li, Feng and Zhang, Hao and Zhang, Kaichen and Li, Yanwei and Liu, Ziwei and Li, Chunyuan , journal =

  32. [32]

    2025 , howpublished =

  33. [33]

    arXiv preprint arXiv:2507.05201 , year =

    Sellergren, Andrew and Kazemzadeh, Sahar and Jaroensri, Tiam and Kiraly, Atilla and Traverse, Madeleine and Kohlberger, Timo and Xu, Shawn and Jamil, Fayaz and Hughes, C. arXiv preprint arXiv:2507.05201 , year =

  34. [34]

    Li, Chunyuan and Wong, Cliff and Zhang, Sheng and Usuyama, Naoto and Liu, Haotian and Yang, Jianwei and Naumann, Tristan and Poon, Hoifung and Gao, Jianfeng , booktitle =

  35. [35]

    and Naumann, Tristan and Wang, Sheng and Poon, Hoifung , journal =

    Zhang, Sheng and Xu, Yanbo and Usuyama, Naoto and Xu, Hanwen and Bagga, Jaspreet and Tinn, Robert and Preston, Sam and Rao, Rajesh and Wei, Mu and Valluri, Naveen and Wong, Cliff and Tupini, Andrea and Wang, Yu and Mazzola, Matt and Shukla, Swadheen and Liden, Lars and Gao, Jianfeng and Crabtree, Angela and Piening, Brian and Bifulco, Carlo and Lungren, M...

  36. [36]

    Proceedings of the 38th International Conference on Machine Learning (ICML) , volume =

    Learning Transferable Visual Models from Natural Language Supervision , author =. Proceedings of the 38th International Conference on Machine Learning (ICML) , volume =. 2021 , publisher =