NTH

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

AuthorsZiqing Yu, Yuhui Tao, Jiayu Huo, Lei Pan, Zilong Xiao, Juecheng Chen, Xiao Li, Jianxuan Li, You Zhou, Zhixing Li, Cong Wang, Beijian Zhang, Chen Chen, Hongyang Lu, Konstantinos Patlatzoglou, Daniel B. Kramer, Jonathan W. Waks, Yangang Su, Fu Siong Ng, Shuo Wang, Yixiu Liang, Junbo Ge

June 1, 2026 2 min read
Watch on YouTube
The one-line take

A CLIP-like foundation model for ECGs learns from millions of ECG-report pairs and can help detect not only common rhythm problems but also rare and subtle cardiovascular diseases.

Key results

2837962
Pretraining ECG-text pairs

ECGCLIP was pre-trained on 2,837,962 ECG studies paired with expert-generated diagnostic reports

1324856
Pretraining patients

Those pretraining pairs came from 1,324,856 patients

45
ECG tasks

The first tier of evaluation included 45 ECG diagnostic tasks

39
ECHO tasks

The second tier of evaluation included 39 echocardiographic tasks

5
Rare disease tasks

The third tier of evaluation included 5 rare cardiac disease tasks

0.900
AF PRAUC

ECGCLIP-R34 achieved a PRAUC of 0.900 for atrial fibrillation on the internal test set

What the paper found

Researchers from Fudan University, Imperial College London, Harvard Medical School, the University of Cambridge, and Medtronic developed ECGCLIP, a signal-language foundation model for cardiovascular assessment that aligns raw 8-lead ECG waveforms with clinician-authored diagnostic reports using CLIP-style contrastive learning. Pretrained on 2,837,962 ECG–text pairs from 1,324,856 patients and validated on nine external cohorts spanning about 1.5 million ECGs, the model was evaluated across 89 tasks: 45 ECG diagnoses, 39 echocardiographic phenotypes, and 5 rare diseases. ECGCLIP-R34 was the strongest variant, improving PRAUC over the Merl-R18 baseline on the internal Zhongshan test set from 0.370 to 0.409 for ECG tasks and from 0.273 to 0.292 for echo tasks, while rare-disease PRAUC rose from 0.011 to 0.045. On atrial fibrillation it reached PRAUC 0.900, on ventricular premature beats 0.913, and on STEMI 0.383, roughly doubling baseline performance for infarction detection. It also enabled opportunistic screening of structural disease from ECG alone, including cardiac amyloidosis with PRAUC 0.201 versus 0.049 for Merl-R18 and Ebstein anomaly with PRAUC 0.253 versus near-zero baselines. A key finding is data efficiency: ECGCLIP-R34 matched or exceeded fully trained baselines using only 10 percent of downstream data. Integrated Gradients and t-SNE showed that the model learned clinically coherent morphology, such as ST elevation, flutter waves, bundle-branch block patterns, and P- and T-wave abnormalities linked to amyloidosis.

Original abstract

Electrocardiography (ECG) is central to cardiovascular care, but conventional AI models are often restricted to common arrhythmias and may generalize poorly across populations or clinically subtle diseases. We developed ECG Contrastive Language-Image Pre-training (ECGCLIP), a signal-language contrastive learning framework that aligns ECG waveforms with expert diagnostic reports. ECGCLIP was pre-trained on 2,837,962 ECG studies from 1,324,856 patients and evaluated on a held-out internal test set plus nine independent external cohorts comprising about 1.5 million ECGs. Evaluation covered 89 downstream tasks, including 45 ECG diagnoses, 39 echocardiographic targets, and 5 rare cardiac diseases, using PRAUC as the primary metric. ECGCLIP consistently improved performance over random initialization and Merl-R18 baselines. On the internal test set, ECGCLIP-R34 achieved strong performance for atrial fibrillation (PRAUC 0.900) and ST-segment elevation myocardial infarction (PRAUC 0.383), with robust generalization across all external cohorts. It also improved low-prevalence and diagnostically elusive diseases, including Ebstein anomaly, constrictive pericarditis, dextrocardia, and cardiac amyloidosis, with internal PRAUC values of 0.253, 0.175, 0.121, and 0.201, respectively. ECGCLIP was data efficient, matching or exceeding full-dataset baseline performance with only 10% of training data. Feature visualization and saliency analysis suggested clinically meaningful representations aligned with established electrocardiographic criteria. These findings indicate that large-scale ECG-report contrastive pre-training can expand routine ECG interpretation beyond common arrhythmias toward broad cardiovascular assessment and opportunistic screening of echocardiographic and rare conditions.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis