AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting

arXiv:2608.09959 · physics.ao-ph, cs.LG · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting".

Jane: The paper was written by Anna Allen, Wessel P. Bruinsma, Michael Maier-Gerber, Harrison Cook, Matthew Chantry et al. from University of Cambridge and European Centre for Medium-Range Weather Forecasts.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making waves in the weather forecasting world, and it's called "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting." Jane, I have to say, just reading that title gets me excited.

Jane: Tom, it should, because this is one of those papers that makes you question how much complexity you actually need. The team took ECMWF's open-source AI weather model, AIFS-Single, and added a relatively simple post-processing step. And that combination matches the performance of Google's purpose-built FNV3 model and the National Hurricane Center's official forecasts.

Tom: And that's the headline, right? A simple correction beating or matching systems that took massive teams and budgets to build.

Jane: Exactly. And I love that the title says "simple correction" because that's genuinely what it is. They're not building a new foundation model from scratch. They're taking an existing one and fixing its biggest weakness, which is that AI weather models tend to underestimate tropical cyclone intensity.

Tom: So the raw AIFS model, it looks at a hurricane and says, "eh, it's a tropical storm," when it's actually a Category five monster.

Jane: That's the failure mode, yeah. The paper shows raw AIFS has a bias of nearly negative twenty-nine knots globally. That means it's systematically under-predicting wind speeds by almost thirty knots on average.

Tom: That's huge. And the correction brings that down to about negative two knots. That's the difference between telling people to board up their windows and telling them to evacuate.

Jane: Right. And I should mention, the authors are from Cambridge and ECMWF, and they're not just reporting results. They're showing that this entire system was designed and built by an AI coding agent, Claude Fable five directed by a single scientist through natural language prompts.

Tom: That's the part that blows my mind. A language model designed the feature engineering, the model architecture, the training pipeline, all of it. In a few hours.

Jane: And that's going to be a theme we come back to, because it changes who can do this kind of work. But first, let's get into what the paper actually shows in terms of results, because the numbers are genuinely impressive.

Summary: Tom: So we're still on "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting," and Jane, you were about to walk us through the actual results.

Jane: Right. So the test set is the entire two thousand twenty-five hurricane season, held out during development. Seventy storms globally. And the metric is mean absolute error for maximum wind speed, averaged over lead times from twelve hours to seven days.

Tom: And the raw AIFS model scores what, around twenty-nine knots globally?

Jane: twenty-eight point nine knots. AIFS-TC cuts that to ten point eight knots. And here's the kicker, FNV3, Google's dedicated tropical cyclone model, scores ten point two knots. Statistically, those are indistinguishable.

Tom: So a simple correction to an open model matches a bespoke model from one of the biggest tech companies in the world.

Jane: And it's not just wind speed. For minimum central pressure, AIFS-TC gets seven point nine millibars of error versus FNV3's seven point eight. Again, statistically tied.

Tom: And what about the official forecasts from the National Hurricane Center?

Jane: In the North Atlantic and East Pacific, where those official forecasts exist, AIFS-TC gets eleven point zero knots, FNV3 gets ten point seven, and the human forecasters get ten point seven. All three are within the confidence intervals of each other.

Tom: So the humans are still in the game, but the gap is essentially zero now.

Jane: Exactly. And here's where it gets really interesting. The paper breaks down performance by storm category, from tropical depression all the way to Category five. And AIFS-TC is competitive at every single level.

Tom: Even Category five? Because those are the most dangerous and the hardest to forecast.

Jane: Even Category five. The error bars are wide because there are only two Category five storms in the test set, but the point estimates are right there with FNV3 and OFCL.

Tom: And I know the paper also looks at rapid intensification, which is when a storm jumps thirty knots or more in twenty-four hours. That's the nightmare scenario for forecasters.

Jane: That's the segment we're about to get into, because that's where raw AIFS completely falls apart. But I want to bring in Lu here, because I think there's a deeper point about what this means for the field.

Lu: Thanks, Jane. I think the remarkable thing is that this isn't a new physical model or a new training paradigm. It's a correction layer. The authors are saying, "the base model has systematic errors, let's learn those errors and subtract them." That's a very old idea in statistics, but applying it to a modern AI weather model at this scale is new.

Tom: And it works because the errors are systematic. The model isn't failing randomly, it's failing predictably.

Lu: Exactly. And that's why a gradient-boosted tree ensemble and a convolutional neural network can learn the correction. They're capturing the pattern in the error.

Jane: And I should mention, the architecture is a blend. Sixty percent weight on the tree model, forty percent on the neural network. The trees handle the derived features, the CNN looks at the three dee atmospheric fields around the storm.

Tom: So it's using both hand-crafted features and raw field data. That's a nice combination.

Lu: It is. And it shows that you don't need one magic model. You need complementary models that see the problem from different angles.

Jane: And that's the hook for our next segment, because the paper also has a lot to say about rapid intensification, which is where the real value of this system shows up.

Improvements: Tom: Back on "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting." Jane, you teased rapid intensification, and I want to hear those numbers.

Jane: So rapid intensification is defined as a thirty knot increase in wind speed over twenty-four hours. Raw AIFS scores a mean absolute error of seventy-one point one knots on those cases. That's not a forecast, that's a guess.

Tom: seventy-one knots of error. The storm is intensifying and the model just sits there saying it's a tropical storm.

Jane: AIFS-TC brings that down to twenty point nine knots. FNV3 gets twenty-three point eight. The official forecasts get twenty point one. So AIFS-TC is statistically tied with both the Google model and the human forecasters.

Tom: And that's the case where lives are actually saved or lost. If you can predict rapid intensification, you can evacuate people before the storm explodes.

Jane: And the paper does something clever in training. They up-weighted the loss for rapid intensification cases by a factor of two, so the model is explicitly pushed to get those cases right.

Tom: So they're telling the model, "these matter more, pay attention."

Jane: Exactly. And it works. But I want to bring in Meng here, because there's a practical question about whether this system can actually run in real time.

Meng: Thanks, Jane. So the paper addresses that directly. They test with operational initial conditions, which means using the real-time storm position and intensity estimates rather than the post-storm best track data that's only available later.

Tom: And does that hurt performance?

Meng: Almost not at all. The mean absolute error goes up by zero point one to zero point three knots depending on the subset. That's well within the confidence intervals. So yes, this can run operationally.

Jane: And what about the compute cost? Because that's the other practical question.

Meng: The correction models are tiny. A gradient-boosted tree and a small CNN. The heavy lifting is done by AIFS, which is already running operationally at ECMWF. So the marginal cost of adding this correction is negligible.

Tom: So any weather service that has access to AIFS output can add this on top.

Meng: And the code is on GitHub. It's open source. That's the part that gets me excited, because it means smaller meteorological agencies, universities, even well-equipped hobbyists could run this.

Lu: And that's the democratization angle. You don't need a team of fifty engineers and a multi-million dollar compute budget. You need a scientist who can describe the problem and an AI agent that can write the code.

Jane: And that's actually the most surprising part of the paper for me. The entire system was built by Claude Fable five an LLM coding agent, directed by a single domain scientist through natural language prompts.

Tom: And the paper says the agent had a tendency to overcomplicate things, and the scientist had to push it to simplify.

Jane: Which is a great reminder that human judgment still matters. The scientist set the goal, defined the evaluation, and validated the outputs.

Meng: And the simplification didn't hurt performance. The final model is just a blend of two model types, five checkpoints each, averaged together.

Lu: I think that's the real lesson. The frontier isn't about building bigger models anymore. It's about knowing where the existing models fail and fixing those specific weaknesses.

Tom: And that's a perfect setup for our conclusion, because I think this paper has implications that go way beyond hurricanes.

Conclusion: Tom: So let's wrap up our discussion of "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting." Jane, give us the one-sentence version.

Jane: A simple, open-source correction to an existing AI weather model matches the performance of the best dedicated tropical cyclone forecasting systems in the world, including human forecasters, and it was built by an AI agent in a few hours.

Tom: And the implications go beyond weather. If this approach works for hurricanes, it could work for other high-stakes prediction problems.

Jane: Absolutely. Think about air quality forecasting, flood prediction, wildfire risk. Any domain where you have a good base model but systematic errors.

Lu: And the agentic coding angle is the bigger story. The paper shows that a single domain scientist, working with an LLM, can build a state-of-the-art system. That's going to change how research gets done.

Meng: And from an operational standpoint, the fact that it works with real-time data and runs at negligible cost means it can be deployed immediately. Not in five years, now.

Jane: And I want to give credit to the authors for being transparent about the limitations. They only trained on two thousand sixteen to two thousand twenty-four they used a single deterministic model, and they didn't include satellite imagery or ocean heat content data, which are known to help.

Tom: So there's room to improve. But the fact that they're already at the frontier without those things is remarkable.

Jane: And the future work section mentions using the full AIFS ensemble for probabilistic forecasts, which would be a natural next step.

Tom: Well, this has been a fantastic discussion. The paper is "AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting," and I think it's going to be cited for years as the example of how to do more with less.

Jane: And on that note, we're ready to move on to the next paper. Thanks for listening, everyone.

Tom: See you next time.

Anna Allen, Wessel P. Bruinsma, Michael Maier-Gerber, Harrison Cook, Matthew Chantry, Richard E. Turner

University of Cambridge · European Centre for Medium-Range Weather Forecasts

physics.ao-ph, cs.LG

Submitted: 2026-08-13

Comments: 6 pages, 5 figures, 2 tables

Code: https://github.com/cambridge-mlg/aifs-tc

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

The gist: The paper presents AIFS-TC, a simple correction to the ECMWF AIFS-Single weather model that achieves performance competitive with the operational state-of-the-art for tropical cyclone (TC) intensity

Key concepts

AIFS-TC
This is a simple correction layer added to the AIFS-Single AI weather model. Its purpose is to fix the raw model's weakness, which is underestimating tropical cyclone intensity, making it competitive with operational forecasting systems.
Systematic Error Correction
The paper shows that raw AI models have systematic errors—they fail in a predictable way. AIFS-TC works by learning these specific errors and subtracting them, which is an application of a statistical idea applied to modern AI weather models.
Rapid Intensification
This refers to a storm increasing its wind speed by thirty knots or more in twenty-four hours. AIFS-TC significantly improves forecasts for rapid intensification cases, bringing its error rate down to levels comparable with dedicated models and human forecasters.
Agentic Coding
The system was built using an AI coding agent, Claude Fable five, directed by a single scientist through natural language prompts. This shows that a domain scientist can use an LLM to build complex features, model architecture, and training pipelines quickly.

Terminology

Summary

The paper presents AIFS-TC, a simple correction to the ECMWF AIFS-Single weather model that achieves performance competitive with the operational state-of-the-art for tropical cyclone (TC) intensity forecasting. The authors state: "We show that a freely available open weather model, ECMWF's AIFS-Single (referred to hereon as AIFS; Lang et al., 2024), combined with a deliberately simple and low-cost learned post-processing step, reaches the level of the operational state-of-the-art currently set by both FNV3 and the National Hurricane Center (NHC) official forecast."

The approach learns an additive correction to AIFS intensity forecasts by modelling AIFS intensity forecast errors. The additive correction is produced from the full AIFS forecast, including "forecasted three-dimensional fields (winds, temperature, geopotential, and specific humidity) cropped to patches centred on the storm and predictions for track and intensity (which are produced from a tracker, see Magnusson et al. (2021), applied to the AIFS)." The model is trained on TCs and AIFS hindcasts for years 2016–2024, with 2025 held out as the test set.

The architecture combines two model types: averaging the output of five gradient boosting machines (GBMs), averaging the output of five convolutional neural networks (CNNs), and combining these two averages with weights 0.6 and 0.4 respectively. The models are separately trained for maximum wind speed and minimum central pressure, resulting in v, p × GBM, CNN × 1,..., 5 = 20 checkpoints in total. The five checkpoints come from five-fold cross-validation. During training, the loss for rapid-intensification cases was up-weighted by a factor two to reflect their practical importance. The inputs include 112 derived features for the GBMs, with the CNNs using 33 of these 112 features and 3D patches of AIFS forecast fields centred on the storm with nz = 4 by selecting pressure levels 850 hPa, 700 hPa, 500 hPa, and 200 hPa. The final maximum wind speed forecast is clipped to [5 kt, 200 kt] and minimum central pressure to [850 mb, 1025 mb].

Key results on the held-out year 2025: Globally, across 70 storms in 2025, AIFS achieves an MAE of 28.9 kt. AIFS-TC is able to reduce this to 10.8 kt, which is comparable to FNV3's 10.2 kt (not statistically significantly different). On the North Atlantic and East Pacific, all three systems perform similarly: 11.0 kt for AIFS-TC, 10.7 kt for FNV3, and 10.7 kt for OFCL (pairwise not statistically significantly different). For minimum central pressure, AIFS achieves an MAE of 11.6 mb globally. AIFS-TC reduces this to 7.9 mb, which is comparable to FNV3's 7.8 mb (not statistically significantly different).

The bias correction is substantial: "raw AIFS forecasts blur the inner core and consequently severely underestimate intensity, leading to a bias of −28.8 kt globally. AIFS-TC reduces this bias to −1.9 kt, which is slightly smaller in magnitude than FNV3's −4.2 kt (statistically significant). For pressure, AIFS has a bias of 10.2 mb. AIFS-TC reduces this to 1.5 mb, which is slightly smaller in magnitude than FNV3's 3.4 mb (statistically significant)."

For rapid intensification (RI), defined as an increase of at least 30 kt in maximum wind speed over 24 h, raw AIFS completely fail[s] to capture RI, resulting in an MAE of 71.1 kt. AIFS-TC brings the MAE down to 20.9 kt, slightly lower than FNV3's 23.8 kt (not statistically significant) and comparable to OFCL's 20.1 kt (not statistically significantly different). For pressure, RI results are 18.8 mb for AIFS-TC versus 20.8 mb for FNV3 (not statistically significantly different). The authors note All systems, including human forecasters (OFCL), retain a substantial negative bias at RI.

Performance is maintained with operational initial conditions: "replacing IBTrACS initial conditions with operationally available ones (ATCF a-deck CARQ rows) increases the MAE in the North Atlantic and East Pacific and in cases of rapid intensification by at most 0.3 kt, well within the 95% confidence intervals."

The paper notes a notable feature: "AIFS-TC was produced by an LLM coding agent (Claude Fable 5) under the direction of a single domain scientist. The scientist specified the goal and evaluation protocol and rigorously verified the outputs; the agent implemented and trained the models and ran the evaluation. The authors state the entire system was autonomously designed and built by a large language model (Claude Fable 5) in a few hours, directed through a small number of natural-language prompts by a single domain scientist. They add There was a general tendency for Claude Fable 5 to find solutions that were overcomplicated. The agent was then directed to simplify its first solution through further interactions. We found that the simplification did not compromise performance."

The authors conclude: "AIFS-TC is a simple correction to AIFS that is competitive with the operational frontier of TC intensity forecasting. On the held-out year 2025, the MAE and bias of AIFS-TC for maximum wind speed and minimum central pressure are generally comparable to and not statistically significantly different from those of FNV3 and, where available, OFCL. This holds even for cases of rapid intensification, which is the regime where AI weather models tend to fail most severely. AIFS-TC therefore demonstrates that state-of-the-art TC intensity forecasting does not require a purpose-built system, but can currently be achieved by a simple machine-learning approach on top of an open AI weather model."

Potential improvements are noted: "Extending training to AIFS hindcasts further back; incorporating the full AIFS ensemble to produce probabilistic intensity forecasts; and augmenting the inputs to the GBM and CNN with observational data that has established intensity skill, such as satellite imagery and ocean heat content (DeMaria and Kaplan, 1994), could all further improve performance."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

1. Hybrid Model Architecture with Dual-Path Correction

  • Implement a 0.6/0.4 weighted blend of gradient-boosted trees (GBM) and convolutional neural networks (CNN) for additive intensity correction, as described in Figure 2.

  • Train separate models for maximum wind speed and minimum central pressure (20 checkpoints total: v,p × GBM, CNN × 1,...,5).

  • Use five-fold cross-validation with storm-level partitioning to generate ensemble predictions from five checkpoints per model type.

2. Enhanced Feature Engineering (112 Features)

  • Derive features from AIFS forecast fields (winds, temperature, geopotential, specific humidity) cropped to 32×32 patches centered on the storm.

  • Include track and intensity predictions from the AIFS tracker.

  • Select four pressure levels (850 hPa, 700 hPa, 500 hPa, 200 hPa) for 3D patches.

  • Incorporate operational best-track data (ATCF b-deck) for real-time deployment.

3. Rapid Intensification (RI) Handling

  • Up-weight the loss function by a factor of 2 for RI cases (≥30 kt increase in 24h) during training.

  • Apply clipping to final forecasts: wind speed [5 kt, 200 kt], pressure [850 mb, 1025 mb].

  • This addresses the critical failure mode where raw AIFS underestimates RI by 71 kt MAE.

4. Operational Readiness

  • Support both IBTrACS and operational ATCF a-deck CARQ initial conditions (performance degradation ≤0.3 kt).

  • Enable forecasts within 3 hours of synoptic time (early model capability).

  • Provide deterministic single-model output (AIFS-Single) for simplicity and low computational cost.

1. State-of-the-Art Intensity Forecasting

  • Achieve MAE of 10.8 kt for maximum wind speed globally (vs. 28.9 kt for raw AIFS), matching FNV3 (10.2 kt) and OFCL (10.7 kt) within statistical significance.

  • Reduce minimum central pressure MAE to 7.9 mb globally (vs. 11.6 mb for AIFS), comparable to FNV3 (7.8 mb).

  • Reduce bias from-28.8 kt (AIFS) to-1.9 kt, outperforming FNV3's-4.2 kt bias.

2. Rapid Intensification Prediction

  • Reduce RI MAE from 71.1 kt (raw AIFS) to 20.9 kt, comparable to OFCL (20.1 kt) and slightly better than FNV3 (23.8 kt).

  • Maintain performance for pressure: 18.8 mb vs. FNV3's 20.8 mb.

3. Basin-Specific Performance

  • Achieve competitive MAE across all basins: North Atlantic (10.8 kt), East Pacific (6.7 kt), West Pacific (7.0 kt), North Indian (2.9 kt), Southern Hemisphere (9.1 kt).

  • Maintain low bias across basins, with the largest remaining bias in Southern Hemisphere (-11.0 kt).

4. Lead-Time Robustness

  • Provide accurate forecasts from 12 to 168 hours with consistent performance.

  • Slightly lower MAE at short lead times (12-48h) compared to FNV3 and OFCL.

5. Operational Deployment

  • Run in real-time using only AIFS forecasts and ATCF a-deck data (no post-storm analysis required).

  • Provide forecasts within 3 hours of initialization, suitable for emergency response.

  • Require minimal computational resources (single deterministic model + simple post-processing).

6. Category-Specific Accuracy

  • Maintain comparable MAE across all intensity categories (TD through Category 5), with no statistically significant degradation at any category.

This system directly addresses the paper's key finding: that a simple, low-cost post-processing approach can match the operational frontier without bespoke models or massive compute budgets.

Sources

Related papers