REVIEW 6 cited by
ECG-Chat: A Large ECG-Language Model for Cardiac Disease Diagnosis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The success of Multimodal Large Language Models (MLLMs) in the medical auxiliary field shows great potential, allowing patients to engage in conversations using physiological signal data. However, general MLLMs perform poorly in cardiac disease diagnosis, particularly in the integration of ECG data analysis and medical report generation, mainly due to the complexity of ECG data analysis and the gap between text and ECG signal modalities. To address these issues, we propose ECG-Chat, a multitask MLLMs focused on ECG medical report generation, providing multimodal conversational capabilities based on cardiology knowledge. We propose a contrastive learning approach that integrates ECG waveform data with text reports, aligning ECG features with reports in a fine-grained manner. This method also results in an ECG encoder that excels in zero-shot report retrieval tasks. Additionally, expanding existing datasets, we constructed a 19k ECG diagnosis dataset and a 25k multi-turn dialogue dataset for training and fine-tuning ECG-Chat, which provides professional diagnostic and conversational capabilities. Furthermore, ECG-Chat can generate comprehensive ECG analysis reports through an automated LaTeX generation pipeline. We established a benchmark for the ECG report generation task and tested our model on multiple baselines. ECG-Chat achieved the best performance in classification, retrieval, and medical report generation tasks. Our code is available at https://github.com/YubaoZhao/ECG-Chat.
Forward citations
Cited by 6 Pith papers
-
ELF: A Family of Encoder-Free ECG-Language Models
A single linear projection from raw ECG to LLM embeddings matches complex encoder-based ECG-language models, while perturbation tests show such models largely ignore the ECG signal.
-
From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
MELP pretrains ECG and text encoders with token-, beat-, and rhythm-level cross-modal supervision and beats prior baselines on several ECG classification benchmarks.
-
SensorLM: Learning the Language of Wearable Sensors
SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.
-
UniECG: Understanding and Generating ECG in One Unified Model
UniECG combines ECG interpretation and text-to-ECG generation in one model by fine-tuning a language model and aligning its output tokens with a pretrained ECG diffusion generator.
-
Signal, Image, or Symbolic: Exploring the Best Input Representation for Electrocardiogram-Language Models Through a Unified Framework
A unified benchmark across six ECG datasets and five text-generation metrics finds tokenized symbolic ECG inputs outperform raw signal and image inputs for ECG-language models.
-
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
Guide-grounded prompting is reported to improve BERTScore of ECG impressions from 0.818 to 0.953, but the supporting tables contain implausible duplicated baseline numbers.
Discussion (0). Continue with ORCID to comment.