A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients

arXiv:2512.18031 · cs.LG, cs.AI · Submitted 2025-12-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We are looking at a heavy hitter today called "A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients."

Jane: That title sounds quite intense, Tom, but it's actually focusing on a very specific and vital part of heart health.

Tom: You mean the Atrial Fibrillation part, which is basically when the heart's upper chambers beat irregularly?

Jane: Exactly, and when that happens in an ICU, it can lead to serious issues like strokes or heart failure.

Lu: I see this as a massive opportunity to build a digital guardian that never sleeps or blinks.

Tom: A digital guardian sounds a bit sci-fi, Lu, but the authors, Sarah Nassar and her team from Queen's University, are making it very real.

Jane: They're trying to use the ECG data that's already being collected from bedside monitors to catch these irregular rhythms automatically.

Meng: My concern is always whether these models can handle the actual chaos of an ICU environment.

Tom: That's a fair point, Meng, because ICU signals are notoriously messy with all the patient movement and equipment noise.

Jane: The researchers are addressing that by providing a specific dataset from a Canadian ICU to help train these systems properly.

Lu: Imagine a world where the monitor doesn't just beep, but actually understands the rhythm's nuance before a human even enters the room.

Meng: If they can make it work despite the noise, it would change how we think about continuous monitoring.

Lalam: This could fundamentally shift the culture of care from reactive emergency response to a state of constant, intelligent vigilance.

Tom: It's a huge shift in how we treat the most vulnerable patients.

Jane: We should look at how they actually put this together in the next part.

Summary: Tom: We've established the goal, so now let's look at how the team actually conducted this research.

Jane: They used two main sources, including their own institutional ICU data and the two thousand twenty-one PhysioNet Challenge dataset.

Tom: I noticed they focused on ten-second segments of ECG data for their testing.

Jane: That's right, and they compared three different ways of using AI to spot the AF.

Meng: I'm curious about those three ways, specifically how they differ in a practical setup.

Tom: They looked at feature-based models, standard deep learning, and these new ECG Foundation Models.

Jane: A Foundation Model is like a brain that's already gone to medical school and just needs a quick refresher on ICU-specific cases.

Lu: It's such a creative leap to take a model trained on millions of signals and adapt it to this niche environment.

Meng: The paper mentions they used transfer learning to get those Foundation Models up to speed.

Tom: And the results were pretty striking, with the ECG-FM hitting an F1 score of zero point eight nine.

Jane: That's a very high score, especially when you consider how difficult the ICU data is to parse.

Lu: Seeing a model perform so well on a small, specialized dataset is a huge win for the field.

Meng: Did they show how the different deep learning architectures handled the one-D signals versus two-D images?

Lalam: They did, and it shows how AI can perceive medical data through different lenses, whether as raw waves or visual plots.

Tom: We'll get into those specific comparisons in just a moment.

Improvements: Tom: Now we're getting into the real meat of the comparison, looking at why some models crushed the others.

Jane: It's fascinating because the standard deep learning models actually struggled when they only had a small amount of ICU data to learn from.

Tom: That's because they were trying to learn everything from scratch, right?

Jane: Exactly, whereas the Foundation Models already had a massive head start.

Meng: I want to talk about the precision aspect, because in an ICU, a false alarm is a huge problem.

Tom: You're talking about alarm fatigue, which can make nurses start ignoring the monitors.

Meng: Precisely, and the paper shows that the ECG-FM achieved a precision of zero point nine eight, which is incredibly high.

Jane: That means it's very unlikely to cry wolf, which is vital for keeping the staff's trust.

Lu: But what if we move beyond just detecting what's happening right now?

Tom: You're thinking about the forecasting part mentioned in the paper, Lu?

Lu: Yes, because if these models can predict an AF episode before it starts, we move into a completely new era of preventative medicine.

Jane: The authors actually suggest that their accurate detection models could be used to label data for future forecasting research.

Meng: That makes a lot of sense from an engineering standpoint, using the current success to build the next generation.

Lalam: This builds a bridge of trust between the clinician and the machine, making the technology a partner rather than a nuisance.

Tom: It really feels like we're standing on the edge of something massive here.

Conclusion: Tom: We've covered a lot of ground with "A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients."

Jane: It's a brilliant piece of work that provides both a new dataset and a roadmap for much better detection.

Lu: I'm just so excited to see how these Foundation Models will evolve to predict heart rhythms before they even change.

Meng: My takeaway is that we finally have a benchmark that respects the messy reality of the ICU.

Lalam: This work will help weave intelligence into the very fabric of patient safety and hospital culture.

Tom: Thanks to everyone for joining us to break this down.

Jane: We'll see you next time for the next big paper on arXiv.

Tom: Goodbye, everyone!

cs.LG, cs.AI

Submitted: 2025-12-19

Updated: 2026-09-11

Comments: 14 pages, 9 figures, 6 tables

Journal ref: IEEE Transactions on Biomedical Engineering (TBME) 2026

DOI: 10.1109/TBME.2026.3715145

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 84/100

The gist: This paper presents a new labelled ICU dataset and establishes performance benchmarks for detecting Atrial Fibrillation (AF) using various artificial intelligence approaches.

Key concepts

Atrial Fibrillation
A condition where the heart's upper chambers beat irregularly. In an ICU setting, this can lead to serious complications like strokes or heart failure, making automatic detection through ECG data vital for patient safety and continuous monitoring.
ECG Foundation Models
These are AI models pre-trained on millions of signals that use transfer learning to adapt to specific tasks. They excel at identifying irregular rhythms in noisy ICU environments, outperforming standard deep learning models that attempt to learn from scratch using limited data.
Alarm Fatigue
This occurs when medical staff begin ignoring monitors due to frequent false alarms. The research emphasizes using high-precision AI models to prevent this, ensuring that alerts are accurate and maintain the trust of clinical staff in a high-stakes environment.

Terminology

Summary

This paper presents a new labelled ICU dataset and establishes performance benchmarks for detecting Atrial Fibrillation (AF) using various artificial intelligence approaches. It is significant because it addresses the scarcity of AF detection studies in the ICU context and provides a foundation for developing automatic patient monitoring system[s] to enable timely treatment in high-risk patients.

Research Objectives and Datasets

The study aims to address gaps in current literature by performing a comprehensive benchmark of AF detection performance across three distinct AI approaches. To achieve this, the researchers utilized two primary data sources:

  • An institutional ICU dataset from a tertiary hospital in Kingston, Ontario, containing almost 600 labelled 10 s four-lead ECGs from unique patients.

  • The large, publicly available 2021 PhysioNet/Computing in Cardiology Challenge dataset to evaluate model performance on a broader scale.

The ICU data is particularly valuable as it provides labels for sinus rhythm and AF (combined with atrial flutter) provided by critical care physicians, which is essential for training models in a unique environment with noisy signals.

AI Methodologies and Approaches

The researchers compared three main data-driven AI approaches to determine which is best suited for the ICU environment. These include:

  1. Feature-based classifiers: Classical machine learning using hand-crafted features such as time domain features, spatial features, and heart rate variability (HRV) features.

  2. Deep learning (DL): This includes 1-D CNNs processing raw signals, 2-D CNNs using ECG images, and models utilizing signal-to-image transformations like spectrograms, scaleograms, and recurrence plots.

  3. ECG foundation models (FMs): Emerging large deep neural networks like ECG-FM and ECGFounder that are pretrained on large volumes of data and can be adapted with minimal additional training.

Experimental Framework

To rigorously evaluate these models, the study implemented four distinct training configurations:

  • Zero-Shot Inference (n=0): Testing pretrained ECG FMs without any additional training.

  • Train on ICU Data (n=298): Training models from scratch using only the small, ICU-specific dataset.

  • Train on PhysioNet Data (n=70,363): Utilizing a large but non-ICU-specific dataset to see if it improves performance.

  • Transfer Learning (n=70,363+298): A strategy where models are first trained on the large PhysioNet data and then further fine-tuned with the ICU training set to tackle the problem of not having a sufficient number of annotated in-house training samples.

Key Findings and Results

The results demonstrate that ECG FMs performed best, followed by DL, then feature-based classifiers. Specifically, the model achieving the top F1 score of 0.89 on the ICU test set was an ECG-FM utilizing a transfer learning strategy. While DL models often struggle with small datasets—where feature-based models and fine-tuned FMs proved more robust to the small training data size—the use of large datasets allowed DL models to eventually surpass feature-based approaches.

The study concludes that these findings demonstrate promising potential for using AI to build an automatic patient monitoring system and provide a way to generate non-expert weak labels for future research into AF forecasting.

Improvements for AI systems

1. Implementation of a Transfer Learning Pipeline using ECG Foundation Models (FMs)

  • Improvement: Transition from training Deep Learning (DL) models or classical ML from scratch to a transfer learning framework that utilizes pre-trained ECG Foundation Models (specifically Transformer-based architectures like ECG-FM). The system must first be trained/fine-tuned on large, multi-site datasets (e.g., 2021 PhysioNet Challenge) and subsequently fine-tuned on domain-specific, high-fidelity ICU telemetry data.

  • Capability: This system can achieve superior F1 scores (targeting >0.89) for Atrial Fibrillation (AF) detection even when labeled ICU-specific data is scarce, significantly outperforming standard 1D or 2D CNNs trained on small local cohorts.

2. Hybrid Multi-Modal Input Processing (Temporal + Morphological)

  • Improvement: Design an ensemble or hybrid architecture that simultaneously processes raw 10-second 1D ECG signals through long-window CNNs and applies signal-to-image transformations—specifically recurrence plots—processed via a 2D Inception v3 or ResNet backbone.

  • Capability: The system can capture both fine-grained temporal dependencies (R-peak irregularities) and complex morphological patterns in the recurrent state, providing higher robustness to the noisy signals and rapid health deteriorations characteristic of the ICU environment.

3. Uncertainty-Aware Classification for Alarm Fatigue Mitigation

  • Improvement: Integrate model calibration and uncertainty estimation layers (e.g., Bayesian neural networks or temperature scaling) into the fine-tuned ECG-FM output. Instead of a binary classification, the system will output a calibrated probability score representing classification confidence.

  • Capability: The system can minimize alarm fatigue in critical care clinicians by suppressing low-confidence detections and only triggering high-priority alerts when the probability of AF exceeds a mathematically optimized threshold, thereby increasing clinical trust and precision.

4. Automated Weak-Labeling Pipeline for AF Forecasting

  • Improvement: Develop a two-stage hierarchical AI system where a high-performing, fine-tuned ECG detector acts as an automated annotator for long-term, continuous ECG recordings to generate weak labels. These labels are then used to train a secondary temporal model (e.g., a Transformer or LSTM) on the identified AF onsets.

  • Capability: This transforms the system from a reactive detection tool into a proactive forecasting engine capable of predicting the onset of an AF episode before it occurs, allowing for timely preventive medical intervention.

5. Lead-Agnostic/Variable-Lead Input Adaptation

  • Improvement: Implement an input layer architecture within the Foundation Model that utilizes zero-shot or fine-tuned mapping to handle variable lead configurations (e.g., 4-lead ICU telemetry vs. 12-lead standard ECG), including strategies for handling missing or unmeasured leads by setting them to zero or utilizing single-lead averaging.

  • Capability: The system can maintain high detection accuracy in real-time bedside monitoring scenarios where signal quality may be degraded or specific ECG leads are disconnected during patient movement.

Abstract

Objective: Atrial fibrillation (AF) is the most common cardiac arrhythmia experienced by intensive care unit (ICU) patients and can cause adverse health effects. In this study, we publish a labelled ICU dataset and benchmarks for AF detection. Methods: We compared machine learning models across three data-driven artificial intelligence (AI) approaches: feature-based classifiers, deep learning (DL), and ECG foundation models (FMs). This comparison addresses a critical gap in the literature and aims to pinpoint which AI approach is best for accurate AF detection. Electrocardiograms (ECGs) from a Canadian ICU and the 2021 PhysioNet/Computing in Cardiology Challenge were used to conduct the experiments. Multiple training configurations were tested, ranging from zero-shot inference to transfer learning. Results: On average and across both datasets, ECG FMs performed best, followed by DL, then feature-based classifiers. The model that achieved the top F1 score on our ICU test set was ECG-FM through a transfer learning strategy (F1=0.89). Conclusion: This study demonstrates promising potential for using AI to build an automatic patient monitoring system. Significance: By publishing our labelled ICU dataset (LinkToBeAdded) and performance benchmarks, this work enables the research community to continue advancing the state-of-the-art in AF detection in the ICU. https://physionet.org/content/kingston-icu-af/

Related papers