Pith. sign in

REVIEW 4 cited by

HeAR -- Health Acoustic Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02522 v1 pith:PZNKKFUJ submitted 2024-03-04 cs.LG cs.AI

classification cs.LGcs.AI
keywords healthacoustichearlearningacousticsaudiodeeptasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Health acoustic sounds such as coughs and breaths are known to contain useful health signals with significant potential for monitoring health and disease, yet are underexplored in the medical machine learning community. The existing deep learning systems for health acoustics are often narrowly trained and evaluated on a single task, which is limited by data and may hinder generalization to other tasks. To mitigate these gaps, we develop HeAR, a scalable self-supervised learning-based deep learning system using masked autoencoders trained on a large dataset of 313 million two-second long audio clips. Through linear probes, we establish HeAR as a state-of-the-art health audio embedding model on a benchmark of 33 health acoustic tasks across 6 datasets. By introducing this work, we hope to enable and accelerate further health acoustics research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A 30-second counting task, embedded with speaker-identification models, predicts male sleep apnea (AUC 0.64) and shows gender- and condition-specific model rankings across a new 7,188-recording clinical speech benchmark.

  2. Adaptable Non-parametric Approach for Speech-based Symptom Assessment: Isolating Private Medical Data in a Retrieval Datastore

    eess.AS 2025-06 conditional novelty 5.0 of 10

    NoNPSA shows that retrieval from a datastore of speech embeddings can assess respiratory symptoms as accurately as fine-tuned self-supervised models, while making data updates and removal easier.

  3. Towards Pre-training an Effective Respiratory Audio Foundation Model

    eess.AS 2025-05 conditional novelty 5.0 of 10

    General audio pre-training (AudioSet) outperforms respiratory-specific pre-training for respiratory sound tasks, and further pre-training on combined AudioSet plus respiratory data yields the best results on the OPERA...

  4. Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Pre-trained acoustic foundation models estimate heart rate from phonocardiogram recordings with accuracy similar to a conventional feature-based method.

Pith tools