Skip to content

Multi-Label 12-Lead ECG Classification on the PTB-XL Dataset: A Comparative Evaluation of Deep Learning Architectures and Heterogeneous Ensemble Approaches

Yunus Emre Mert, Ece Akdoğan, Hüseyin Üvet

eess.SParXiv:2609.12803

Abstract

This study aimed to compare the performance of different deep learning architectures and heterogeneous ensemble learning approaches for multi-label 12-lead ECG classification on the PTB-XL dataset. Five different models, namely 1D-ResNet18, Bidirectional Mamba, xLSTM, CWT-ViT-KAN, and the pre-trained ECGFounder, were evaluated. Utilizing the recommended split structure of the PTB-XL dataset, folds 1-8 were allocated as the training set, fold 9 as the validation set, and fold 10 as the independent test set. In addition to the individual models, three different ensemble approaches were investigated: Equal-Weight Soft Voting, Validation-Weighted Soft Voting, and stacking. Among the individual models, the highest performance was achieved by ECGFounder, with a Macro AUROC of 0.930 and a Macro AUPRC of 0.823. For the ensemble models, the highest values in the primary macro performance metrics were obtained by the stacking approach, achieving a Macro AUROC of 0.936, a Macro AUPRC of 0.836, and a Macro F1 of 0.763. The highest subset accuracy of 0.630 was achieved using the Validation-Weighted Soft Voting method. The findings indicate that heterogeneous ensemble models, which combine different representation learning approaches, can provide additional performance improvements over individual models in multi-label ECG classification.

Create a lesson