AI for pRedicting Exacerbations in KIDs with aSthma (AIRE-KIDS)

arXiv:2511.01018 · cs.AI · Submitted 2025-11-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "AI for pRedicting Exacerbations in KIDs with aSthma (AIRE-KIDS)".

Jane: The paper was written by Hui-Lee Ooi, Nicholas Mitsakakis, Margerie Huet Dastarac, Roger Zemek, Amy C. Plint et al. from CHEO Research Institute and University of Ottawa.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. We've got a paper that's near and dear to a lot of families out there, and it's called "AI for pRedicting Exacerbations in KIDs with aSthma," or AIRE-KIDS for short. Jane, I have to say, the acronym alone makes me smile.

Jane: It's a good one, Tom. And the title actually tells you exactly what they're trying to do. They want to use AI to predict when a kid with asthma is going to have a really bad flare-up that lands them back in the emergency department or even in the hospital.

Tom: Right, and that's such a huge deal. My nephew has asthma, and every time he gets a cold, it's this anxious wait to see if it's going to turn into something serious. This paper is basically trying to give doctors a heads-up.

Jane: Exactly. And the "KIDs" part is important, because kids aren't just little adults. Their bodies react differently, and the way asthma hits them is different. So you can't just take an adult model and apply it to them.

Tom: So the team behind this is from the CHEO Research Institute in Ottawa, and they've got a really interesting mix of people. You've got pediatricians, obviously, but also folks who specialize in data and machine learning.

Jane: And that's the key, right? You need the doctors who understand the disease to tell the data scientists what matters, and you need the data scientists to figure out how to actually make sense of all that messy hospital data.

Tom: Messy is the right word. I mean, they pulled in stuff like how long the kid waited in the ED, what their triage score was, even air quality data from their neighborhood. It's a ton of information.

Jane: And the goal is pretty straightforward, even if the math is complex. They want to build a tool that, when a kid walks into the ER with asthma, can look at all this information and say, "This one's high risk, we should really get them into the comprehensive asthma program."

Tom: So it's not just about treating the emergency. It's about preventing the next one.

Jane: Exactly. And that's the part that gets me excited. This isn't just a lab experiment. They're building it to be used right at the bedside, to help real doctors make real decisions.

Tom: And we're going to get into the nitty-gritty of how well it actually works in a second. But first, I want to know, what was the biggest surprise for you when you first read through this?

Jane: Honestly, Tom, it was the fact that they compared their fancy new models against the old-school rule that the hospital was already using. They didn't just say, "Look, our AI is good." They said, "Look, our AI is better than what you're doing right now." That's the kind of proof that actually changes practice.

Summary: Tom: So we're digging into the AIRE-KIDS paper now, and Jane, you just teased that they compared it to the old way of doing things. Let's talk about what they actually found.

Jane: So the old rule, the "best practice alert," was pretty simple. If a kid came back to the ED within a year and had a certain severity score, they'd flag them for the asthma program. It was a blunt instrument.

Tom: And the new model, the AI, is a lot more sophisticated. They used something called a LightGBM model, which is a type of machine learning that's really good at finding patterns in big tables of data.

Jane: Right. And they trained it on two different groups of patients. One group was from before COVID, and the other was from after, to make sure the model worked in different circumstances.

Tom: And the results? I mean, the headline is that the AI did better, but how much better?

Jane: Well, for predicting repeat ED visits, the AI model got an F1 score of zero point five one, while the old rule only got zero point three three four. And for hospital admissions, it was zero point three seven five versus zero point three one three. The F1 score is a way of measuring both how good the model is at catching the right kids and not crying wolf on the wrong ones.

Tom: So it's a meaningful jump. But it's not perfect, right? I mean, zero point five one is not zero point nine.

Jane: No, it's not perfect. And they're honest about that. But think about what it's replacing. The old rule was basically guessing based on a couple of factors. The new model is looking at things like whether the kid has a food allergy, how complex their medical history is, and even their age.

Tom: That's the part that's wild to me. A food allergy? How does that connect to asthma?

Jane: It's a known clinical link, actually. Kids with food allergies often have more severe asthma. And that's the beauty of the model—it can weigh all these little factors that a busy doctor might not have time to piece together in the moment.

Tom: So it's not just about the big, obvious stuff. It's about the subtle combination of things.

Jane: Exactly. And they also found that the model was better than just saying "everyone is high risk." That sounds silly, but it's a real baseline. If you referred everyone, you'd catch all the high-risk kids, but you'd also overwhelm the asthma program with kids who are fine.

Tom: And that's the whole point. The program has limited resources. They need to target the kids who will benefit the most.

Jane: Right. And that's why the F1 score matters. It's about finding that balance. And the AIRE-KIDS model does that much better than the current approach.

Tom: So, real quick, what was the one thing that made the biggest difference in the model's prediction? Was it the air quality stuff?

Jane: Actually, no. That was a surprise. They thought environmental factors would matter, but they didn't. The biggest predictors were things like a previous asthma ED visit and the kid's medical complexity. It's the history that tells you the most about the future.

Improvements: Tom: So we know the AIRE-KIDS model works better than the old rule. But what does that actually mean for the kids and the hospital? What's the improvement in practice?

Jane: That's the big question, Tom. And the paper gets into this. The whole point is to use this model as a decision-support tool in the emergency department. So when a kid comes in with asthma, the model runs in the background and gives the doctor a risk score.

Meng: And from an engineering standpoint, that's a big deal. They specifically chose a model that's lightweight enough to run on local hospital computers. They didn't want to have to send patient data to the cloud, which would be a privacy nightmare.

Tom: Oh, that's a great point, Meng. So it's not just about accuracy. It's about being practical and respecting privacy rules.

Meng: Exactly. And they also made the model parsimonious. They took all those sixty-eight features they started with and boiled it down to just six for the ED visit model. That makes it way easier to integrate into the existing electronic medical record system.

Jane: And that's the kind of thing that makes a tool actually get used. If it requires a data scientist to run it, it's never going to happen. But if it's just a little alert that pops up, that's something a doctor can act on.

Lu: I'd add that the improvement here is also about shifting the focus from reaction to prevention. The old rule was basically saying, "You've been here before, so you might come back." The new model is saying, "Based on this complex set of signals, this child is at high risk, and here's a chance to intervene."

Tom: So it's not just a better calculator. It's a different philosophy of care.

Lu: Precisely. And that's why the authors are so excited about the potential to reduce morbidity. If you can get a high-risk kid into a comprehensive asthma education program, you can actually change the trajectory of their disease.

Jane: And there's a real-world constraint they had to deal with. The asthma program at CHEO has limited capacity. They can't see every kid. So the model helps them use that capacity where it matters most.

Meng: And that's the practical impact. It's not just about a number on a screen. It's about making sure the right kids get the right care at the right time.

Tom: So, Lu, you're the visionary here. Where does this go from here? Is this just for CHEO, or could this spread?

Lu: Well, the authors note that the model would likely need to be recalibrated for other centers, because different hospitals have different patient populations. But the framework, the approach, is definitely transferable. This is a blueprint for how to build a practical, privacy-preserving clinical AI tool.

Jane: And that's the exciting part. This isn't just a paper that sits on a shelf. It's a working system that they're planning to deploy and evaluate prospectively. That's the gold standard.

Conclusion: Tom: Alright, we've spent a good chunk of time on the AIRE-KIDS paper, and I think we've covered a lot of ground. Jane, can you wrap it up for us?

Jane: Sure, Tom. So the AIRE-KIDS paper is all about using AI to predict which kids with asthma are going to have a severe exacerbation that lands them back in the hospital. They built a model that's more accurate than the current best practice, and they did it using data that's already being collected in the electronic medical record.

Tom: And the key takeaway for me was that it's not just a better model. It's a more practical one. They made it work with fewer variables, they made it run locally to protect privacy, and they compared it against a real-world baseline.

Jane: Exactly. And they showed that the LightGBM approach beat out the newer, fancier large language models. That's a reminder that sometimes the most powerful tool isn't the most glamorous one.

Lu: And the potential impact is significant. If this works in practice, it could mean fewer emergency visits, fewer hospitalizations, and better long-term outcomes for kids with asthma. It's a way to make the healthcare system smarter and more equitable.

Meng: And from a technical perspective, it's a solid piece of engineering. They did the hard work of validation, calibration, and thinking about deployment from day one.

Tom: Well said. So we're saying goodbye to AIRE-KIDS, but we're taking away a lot of hope. It's a great example of how AI can be used responsibly and effectively in medicine.

Jane: And with that, we'll get ready to dive into the next paper. Thanks for listening, everyone. We'll see you next time.

Tom: Take care, folks.

Hui-Lee Ooi, Nicholas Mitsakakis, Margerie Huet Dastarac, Roger Zemek, Amy C. Plint, Jeff Gilchrist, Khaled El Emam, Dhenuka Radhakrishnan

CHEO Research Institute · University of Ottawa

cs.AI

Submitted: 2025-11-02

Updated: 2026-08-18

Code: https://github.com/meta-llama/llama-models

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

The gist: Recurrent exacerbations remain a common yet preventable outcome for many children with asthma.

Key concepts

AIRE-KIDS
This is the name of a study using artificial intelligence to predict severe asthma flare-ups in children. The goal is to provide doctors with an early warning, allowing them to intervene and prevent emergency room visits or hospitalizations.
F1 Score
This metric measures how effective the AI model is at identifying high-risk patients while avoiding false alarms. A higher F1 score indicates a better balance between catching all high-risk cases and not misclassifying healthy children.
LightGBM Model
This is a specific type of machine learning algorithm used in the study. It is highly effective at finding patterns within large datasets, allowing the AI to weigh complex factors like medical history and food allergies to make accurate predictions.
Decision-Support Tool
The AI model functions as a tool for doctors in the emergency department. It runs in the background, providing a risk score based on patient data, helping clinical staff decide which children need immediate intervention.

Terminology

Summary

Recurrent exacerbations remain a common yet preventable outcome for many children with asthma. Machine learning (ML) algorithms using electronic medical records (EMR) could allow accurate identification of children at risk for exacerbations and facilitate referral for preventative comprehensive care to avoid this morbidity. The authors developed ML algorithms to predict repeat severe exacerbations (i.e., asthma-related ED visits or future hospital admissions) for children with a prior asthma ED visit at a tertiary care children's hospital.

Up to 25% of children with an asthma ED visit will have a repeat visit in the following year, despite numerous evidence-based treatments available to optimize asthma control. With current shortages in primary care, preventative intervention is frequently missed in the ambulatory setting. The authors note that there is a lack of guidance or agreed consensus on how to identify this vulnerable group of children who are most at risk for repeat severe asthma exacerbations.

The study used retrospective Epic EMR data to train ML models for predicting future severe asthma exacerbations among patients treated for an asthma exacerbation in the ED at the Children's Hospital of Eastern Ontario (CHEO). A future severe asthma exacerbation was defined as an ED re-visit or hospitalization with asthma as the most responsible diagnosis within one year of the initial asthma ED visit.

Two retrospective datasets were used: a training dataset (pre-COVID-19) accrued patients over 2 years between February 2017 to March 2019, with outcomes observed up to March 2020, and a post-COVID validation dataset with an accrual period from July 2022 to April 2023, with outcomes observed up to April 2024. Patients with an index visit for asthma during the peak COVID19 period (March 2020-June 2022) were excluded due to significantly lower frequency of asthma exacerbations observed following provincial infection control measures.

A total of 68 clinical features were extracted from the EMR and further linked to air pollution and neighborhood marginalization data. Features included:

  • Prior health care utilization: previous asthma-related ED visits or admissions, prior non-asthma respiratory ED visits or admissions within one year prior to the index visit

  • Operational variables: day and month of the index visit, time from arrival to triage or disposition, hourly average number of patients present in the ED, average wait time during each patient's visit

  • Demographic features: age during the index visit, sex, whether the patient has a primary care physician, distance of the hospital from the postal code of the patient's home, province of residence

  • Severity features: first oxygen saturation recording, timing of order and administration of oral corticosteroids, Pediatric Respiratory Assessment Measure (PRAM) score at triage, Canadian Triage Acuity Score (CTAS)

  • Clinical variables: presence of food and environmental allergens, lab test results such as serum eosinophils count and immunoglobulin E levels, medical complexity as reflected in prior visits to subspecialty clinics, presence of specific comorbidities

  • Environmental and social features (for Ontario residents): Air Quality Health Index (AQHI), Nitrogen Dioxide, Ozone, and PM2.5 readings at different intervals (1 day, 2 days, 3 days prior to index visit), and neighborhood marginalization quintile using the Ontario marginalization (ON-Marg) index

Patients were included if they had an asthma diagnosis (ICD code J45, J46) at the index encounter during the accrual period. Patients were excluded if they had attended a respirology appointment or received comprehensive asthma education prior to the index visit, or if they were admitted to hospital during the same encounter as the index asthma ED visit.

Two types of boosted decision trees (LGBM and XGBoost) and three open-source LLMs with fine-tuning (DistilGPT2, Llama 3.2 1B, and Llama-8b-UltraMedical) were evaluated. The hyperparameters of boosted tree models were tuned and optimized using Bayesian optimization with 5-fold cross-validation. Predicted values were calibrated using beta calibration. For LLMs, a textual encoder transformed each tabular record into text, and predictions were generated five times for each patient and averaged. Due to privacy rules, commercial off-site LLMs could not be utilized.

The F1 score was deemed the most appropriate evaluation metric for the clinical context, being the harmonic mean of precision and recall. The authors note that "capacity within the CHEO asthma program is resource-limited and may not be available to all patients with the predicted outcome. Therefore, we wanted a model that maximizes the proportion of patients that were predicted as high risk that were actually high risk (precision), hence ensuring that the resources are used most effectively." Cutoffs were selected to maximize the F1 score. Three baselines were established: (a) the 'CHEO best practice' rule currently used in the ED, (b) a naïve prediction rule predicting everyone is high risk, and (c) a random prediction of risk status based on a 0.5 probability. AUC was also reported to enable comparison with previous studies.

There were 2,716 patients in the training cohort and 1,237 in the validation cohort. The majority were preschoolers (mean age 4.4-4.5 years, median 3 years), 34.3-36.5% were female, and the majority (72.4% - 81.8%) resided in Ontario. Outcome distributions were similar between cohorts: 29.3% repeat asthma ED visits in training and 28.5% in validation cohorts; 11.9% and 17.8% for admissions respectively.

The current CHEO best practice rule for alerting ED clinicians about which patients to refer to the CHEO asthma program achieved F1 scores of 0.270 (ED visit) and 0.248 (admission) on training data, and 0.334 (ED visit) and 0.313 (admission) on validation data.

Internal validation on the training dataset using nested cross-validation showed AUC values ranging from 0.530 to 0.666 for predicting ED visits and 0.5 to 0.74 for predicting admissions, with lower AUCs observed for the LLMs. The best performance was for the LGBM model.

When applied to the validation dataset, F1 scores for predicting ED visits ranged from 0.292 to 0.503, and from 0.140 to 0.342 for predicting admissions. In all cases, F1 scores were highest for the LGBM models and lowest for the LLMs. The LGBM model achieved an AUC of 0.702 with F1 of 0.503 for ED visits, and AUC of 0.673 with F1 of 0.342 for admissions.

SHAP analysis revealed the most predictive features for repeat asthma ED visits included: previous asthma ED visit, CTAS, medical complexity, food allergy, ED visit for other respiratory diagnoses, and age. For predicting future asthma hospitalization, the top features included: medical complexity, previous asthma ED visit, average wait time, PRAM score during the index visit, and food allergy.

Based on SHAP feature importance, parsimonious models were created with reduced feature sets:

  • AIRE-KIDS ED model: features included prior asthma ED visit, CTAS, complexity, food allergy, prior ED visits for non-asthma respiratory diagnoses, and age. This achieved an AUC of 0.712 and F1 of 0.51 at a cutoff threshold of 0.250 on the validation dataset.

  • AIRE-KIDS HOSP model: features included complexity, prior asthma ED visit, average wait time in the ED, PRAM at triage, and food allergy. This achieved an AUC of 0.65 and F1 score of 0.375 at a cutoff threshold of 0.170.

The authors note that this is a nontrivial improvement over the current decision rule which has F1=0.334. The best naïve decision rule that predicts all patients having high risk has an F1 score of 0.442 and 0.302 respectively for ED visits and admissions.

The authors conclude that "AIRE-KIDS ED and AIRE-KIDS HOSP are new ML models for accurately predicting repeat asthma related ED visits or future asthma-related hospitalizations in children, and are superior to current decision-making processes for triaging referrals for comprehensive preventative care from the ED. They further state that LGBM remains a superior ML approach for modeling this type of clinical prediction, outperforming newer LLMs, despite their popularity in other applications, though ongoing research in this area is needed as LLM's continue to evolve."

The next steps involve local deployment and prospective evaluation at our centre in the form of a clinical decision support system and subsequent validation at other pediatric centres. Successful implementation "could redefine how we triage and manage pediatric asthma in the acute care setting, supporting more equitable access to comprehensive or specialist care while reducing the morbidity associated with repeat severe exacerbations."

The authors acknowledge several limitations: potential lack of generalizability to other centres requiring additional validation; restriction to local healthcare visit data potentially missing outcomes at other centres (though more than 94% of children within 50km attend CHEO); and missing variables difficult to obtain from EMR data including adherence to asthma controller medications or exposure to cigarette smoke in the home. Including patients who received asthma education or attended a respirology clinic visit in the outcome variables introduces error, though the authors argue results are arguably better than presented here because of this.

Improvements for AI systems

Based on the AIRE-KIDS paper, here are the specific improvements I can make to AI systems, and what the improved systems can do:


Improvement: The paper demonstrates that LGBM (Light Gradient-Boosting Machine) significantly outperforms three fine-tuned open-source LLMs (DistilGPT2, Llama 3.2 1B, Llama-8b-UltraMedical) for tabular clinical prediction tasks. The AUC for ED visit prediction was 0.666 (LGBM) vs. 0.530–0.587 (LLMs). For hospital admission, LGBM achieved 0.742 AUC vs. 0.500–0.650 for LLMs.

What the improved AI system can do: Automatically default to gradient-boosted tree models (LGBM/XGBoost) for structured EMR data, reserving LLMs only for unstructured text. This avoids the common pitfall of deploying LLMs on tabular data where they underperform, saving computational cost and improving accuracy.

Improvement: The paper used SHAP values to reduce from 68 features to just 6 for ED visit prediction (prior asthma ED visit, CTAS, medical complexity, food allergy, prior non-asthma respiratory ED visits, age) and 5 for hospital admission (medical complexity, prior asthma ED visit, average ED wait time, PRAM score, food allergy). This reduction maintained performance (AUC 0.712 for ED visits) while simplifying deployment.

Improvement: Instead of relying solely on AUC, the paper optimized the decision threshold to maximize the F1 score (harmonic mean of precision and recall). The AIRE-KIDSED model achieved F1 = 0.510 vs. the current best-practice alert's F1 = 0.334, and vs. a naïve predict all high-risk baseline of F1 = 0.442.

Improvement: The paper applied beta calibration to convert LGBM pseudo-probabilities into true probabilities. This is essential because boosted trees produce uncalibrated outputs that can mislead clinicians about actual risk.

Improvement: The models were trained on pre-COVID data (Feb 2017–Feb 2019) and validated on post-COVID data (Jul 2022–Apr 2023), demonstrating robustness despite significant changes in healthcare utilization patterns during the pandemic.

Improvement: The paper accounted for patients who received asthma education or specialist care (which reduces exacerbation risk) by including these as positive outcomes. This prevents underestimating model performance.

Improvement: The model incorporated operational ED features such as average wait time and patient volume, which were predictive of hospital admission. This goes beyond traditional clinical features.

Improvement: The paper emphasized that models trained at one center may require local recalibration due to differences in case mix and practice patterns. They prioritized parsimony and transparency to facilitate transfer.

Improvement: SHAP analysis revealed clinically plausible predictors (e.g., prior ED visits, triage acuity, food allergy), which builds clinician trust and supports adoption.

Improvement: The paper developed separate models for ED revisit (AIRE-KIDSED) and hospital admission (AIRE-KIDSHOSP), recognizing these as distinct outcomes with different clinical implications.

In summary, the improved AI system can:

  • Accurately predict repeat asthma ED visits (AUC 0.712) and hospitalizations (AUC 0.65) in children

  • Outperform current clinical decision rules by 53% (F1: 0.510 vs. 0.334)

  • Operate in real-time at the ED point of care with only 5–6 easily obtainable features

  • Provide calibrated, explainable risk scores that support clinical decision-making

  • Maintain performance across temporal shifts and be adaptable to other centers

  • Optimize resource allocation for limited-capacity preventative care programs

Abstract

Recurrent exacerbations remain a common yet preventable outcome for many children with asthma. Machine learning (ML) algorithms using electronic medical records (EMR) could allow accurate identification of children at risk for exacerbations and facilitate referral for preventative comprehensive care to avoid this morbidity. We developed ML algorithms to predict repeat severe exacerbations (i.e. asthma-related emergency department (ED) visits or future hospital admissions) for children with a prior asthma ED visit at a tertiary care children's hospital. Retrospective pre-COVID19 (Feb 2017 - Feb 2019, N=2716) Epic EMR data from the Children's Hospital of Eastern Ontario (CHEO) linked with environmental pollutant exposure and neighbourhood marginalization information was used to train various ML models. We used boosted trees (LGBM, XGB) and 3 open-source large language model (LLM) approaches (DistilGPT2, Llama 3.2 1B and Llama-8b-UltraMedical). Models were tuned and calibrated then validated in a second retrospective post-COVID19 dataset (Jul 2022 - Apr 2023, N=1237) from CHEO. Models were compared using the area under the curve (AUC) and F1 scores, with SHAP values used to determine the most predictive features. The LGBM ML model performed best with the most predictive features in the final AIRE-KIDS ED model including prior asthma ED visit, the Canadian triage acuity scale, medical complexity, food allergy, prior ED visits for non-asthma respiratory diagnoses, and age for an AUC of 0.712, and F1 score of 0.51. This is a nontrivial improvement over the current decision rule which has F1=0.334. While the most predictive features in the AIRE-KIDS HOSP model included medical complexity, prior asthma ED visit, average wait time in the ED, the pediatric respiratory assessment measure score at triage and food allergy.

Sources

Related papers