REVIEW 7 cited by
SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sleep is a complex physiological process evaluated through various modalities recording electrical brain, cardiac, and respiratory activities. We curate a large polysomnography dataset from over 14,000 participants comprising over 100,000 hours of multi-modal sleep recordings. Leveraging this extensive dataset, we developed SleepFM, the first multi-modal foundation model for sleep analysis. We show that a novel leave-one-out approach for contrastive learning significantly improves downstream task performance compared to representations from standard pairwise contrastive learning. A logistic regression model trained on SleepFM's learned embeddings outperforms an end-to-end trained convolutional neural network (CNN) on sleep stage classification (macro AUROC 0.88 vs 0.72 and macro AUPRC 0.72 vs 0.48) and sleep disordered breathing detection (AUROC 0.85 vs 0.69 and AUPRC 0.77 vs 0.61). Notably, the learned embeddings achieve 48% top-1 average accuracy in retrieving the corresponding recording clips of other modalities from 90,000 candidates. This work demonstrates the value of holistic multi-modal sleep modeling to fully capture the richness of sleep recordings. SleepFM is open source and available at https://github.com/rthapa84/sleepfm-codebase.
Forward citations
Cited by 7 Pith papers
-
STEAM: A Spatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding
STEAM, a dual-branch spatio-temporal mixture-of-experts EEG model with two-stage pre-training, reports the best average rank across seven downstream EEG datasets and fourteen evaluation settings.
-
Agentic AI-enabled discovery across large-scale sleep physiology
A human-guided agentic AI pipeline analyzed ~124,000 polysomnograms and reported five sleep findings, including an association between reduced N2 brain coupling and incident Parkinson's (HR 1.48) and Alzheimer's (HR 1...
-
OSF: On Pre-training and Scaling of Sleep Foundation Models
Channel-masked self-supervised pretraining on a 166,500-hour multi-source sleep corpus yields OSF, which generalizes better to missing channels and scales with data and model size.
-
SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures
SleepMaMi, a dual-encoder sleep foundation model pretrained on 158K hours of PSG, matches or beats existing sleep foundation models on staging, apnea segmentation, and disease prediction.
-
PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
PiCME shows contrastive learning peaks at three modalities in MIMIC, and a Modality-Gated LSTM with contrastively learned weights improves five-modality mortality prediction over supervised baselines.
-
Self-DANA: A Resource-Efficient Channel-Adaptive Self-Supervised Approach for ECG Foundation Models
Self-DANA combines dimension-adaptive pooling with random lead selection to fine-tune ECG foundation models on reduced-lead inputs, cutting memory and time while maintaining diagnostic accuracy.
-
Estimating Markers of Driving Stress through Multimodal Physiological Monitoring
A multimodal physiological classifier distinguishes stressor periods from free driving in a simulator with AUROC 0.812, but the per-second real-time claim is weakened by a centered, non-causal window.
Discussion (0). Sign in to comment.