Pith. sign in

REVIEW 2 cited by

EchoFM: Foundation Model for Generalizable Echocardiogram Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23413 v2 pith:FIE4POCD submitted 2024-10-30 cs.CV

classification cs.CV
keywords echocardiographyechofmfoundationmodelmodelstasksacrossdownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models have recently gained significant attention because of their generalizability and adaptability across multiple tasks and data distributions. Although medical foundation models have emerged, solutions for cardiac imaging, especially echocardiography videos, are still unexplored. In this paper, we introduce EchoFM, a foundation model specifically designed to represent and analyze echocardiography videos. In EchoFM, we propose a self-supervised learning framework that captures both spatial and temporal variability patterns through a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. This framework can effectively capture the spatio-temporal dynamics of echocardiography and learn the representative video features without any labels. We pre-train our model on an extensive dataset comprising over 290,000 echocardiography videos covering 26 scan views across different imaging modes, with up to 20 million frames of images. The pre-trained EchoFM can then be easily adapted and fine-tuned for a variety of downstream tasks, serving as a robust backbone model. Our evaluation was systemically designed for four downstream tasks after the echocardiography examination routine. Experiment results show that EchoFM surpasses state-of-the-art methods, including specialized echocardiography methods, self-supervised pre-training models, and general-purposed pre-trained foundation models, across all downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Representation Learning for Real-Time Ultrasound Analysis

    eess.IV 2025-09 reject novelty 4.0 of 10

    A frame-wise masked autoencoder with a temporal contrastive loss achieves 0.88 AUROC for binary EF classification on EchoNet-Dynamic, below the cited 0.93 AUROC of ECHO-VISION-FM.

  2. QueEn: A Large Language Model for Quechua-English Translation

    cs.CL 2024-12 reject novelty 3.0 of 10

    QueEn reports BLEU 17.6 for Quechua-English translation, but its internal table shows BLEU 0.235, the method is not reproducible, and the translation direction is inconsistent.

Pith tools