AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons

arXiv:2502.16036 · physics.ao-ph, cs.LG · Submitted 2025-02-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons".

Jane: The paper was written by Daosheng Xu, Zebin Lu, Jeremy Cheuk-Hin Leung, Dingchi Zhao, Yi Li et al. from Guangzhou Institute of Tropical and Marine Meteorology and Guangdong Provincial Key Laboratory of Regional Numerical Weather Prediction and China Meteorological Administration and Key Laboratory of Physical Oceanography and Ministry of Education and Institute for Advanced Ocean Study and Frontiers Science Center for Deep Ocean Multispheres and Earth System and College of Oceanic and Atmospheric Sciences and Ocean University of China and College of Meteorology and Oceanography and National University of Defense Technology and Guangdong Meteorological Observatory and College of Atmospheric Science and Lanzhou University and National Meteorological Centre and State Key Laboratory of Tibetan Plateau Earth System and Resources and Environment and Institute of Tibetan Plateau Research and Chinese Academy of Sciences.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the weather forecasting world, and the title really tells you everything you need to know: "AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons."

Jane: Tom, that title is a bit of a gut punch, isn't it? I mean, we've been hearing so much about how AI weather models are catching up to, or even beating, the old-school physics-based models. And now this paper comes along and says, not so fast.

Tom: Exactly. And I love that they're being so specific. They're not saying AI is bad at everything. They're saying AI is worse at one very particular, very dangerous thing: typhoons that suddenly turn.

Jane: Right, and that's such an important distinction. A typhoon that just keeps going in a straight line is relatively easy to predict. But a typhoon that makes a sharp turn, like a car swerving without warning, that's the kind of thing that catches people off guard and causes real disasters.

Tom: And the paper points to a specific example, Typhoon Khanun in two thousand twenty-three. This thing made two sharp turns in five days as it passed through the Ryukyu Islands. And the AI model, Pangu-Weather, just couldn't keep up with the traditional European model.

Jane: So the title is basically saying, when it comes to the rare and the extreme, the old physics-based approach still has the edge. That's a really important reality check for anyone who thinks AI is just going to sweep in and solve everything.

Tom: It really is. And I think the key word in the title is "still." It's not saying AI will never get there. It's saying, right now, in this specific area, the traditional models are still the ones you want to trust.

Jane: And that's a good thing to know, because if you're a forecaster in a country that's about to get hit by a sudden-turning typhoon, you need to know which tool is going to give you the most reliable information.

Tom: Absolutely. So we've got the headline. But what's the actual evidence behind this claim? Let's dig into the details of the paper in the next segment.

Summary: Jane: So we're back, and we're still talking about "AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons." Tom, let's get into what the researchers actually did.

Tom: Okay, so they looked at one hundred four typhoons in the Northwest Pacific from two thousand twenty to two thousand twenty-four. And they split them into three groups: ordinary ones, sudden-turning ones, and looping ones, which are the ones that do a little circle.

Jane: And that's a smart way to break it down, because it lets you see if the AI model's performance changes depending on how weird the typhoon's path is.

Tom: Right. And what they found is that for the ordinary typhoons, the AI model, Pangu-Weather, actually beat the traditional European model, ECMWF-IFS. The average error was about thirteen point five percent smaller. So AI is winning the easy cases.

Jane: And that's consistent with what the original Pangu-Weather paper claimed, right? That it's generally more accurate.

Tom: Exactly. But here's where it gets interesting. When they looked at the sudden-turning typhoons, the tables turned completely. The AI model's errors were about nine point four percent larger than the traditional model's. So on the hard cases, the physics-based model wins.

Jane: So it's not that AI is bad, it's that it's good at the common stuff and bad at the rare stuff. And the rare stuff is exactly what causes the most damage.

Tom: And they even looked at a specific case, Typhoon Khanun, which I mentioned earlier. The AI model's average track error over twenty-four to one hundred twenty hours was eighteen point eight percent worse than the traditional model. And the gap just got bigger as the forecast lead time increased.

Jane: So the longer the forecast, the worse the AI model does relative to the traditional one. That's a big deal, because you need those long lead times to evacuate people.

Tom: And here's the kicker. They also compared the AI model to human forecasters. And for the first couple of days, the AI model was better than the humans. But by day five, the difference basically disappeared. So the AI model's advantage fades away exactly when you need it most.

Jane: That's a really sobering finding. So the summary is, AI is great at the average, but it's not ready for the extreme. And the paper has some ideas about why that is, which we should get into.

Tom: Yeah, let's talk about that in the next segment. What's causing this blind spot?

Improvements: Tom: We're back with "AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons." And Jane, we've established that AI is struggling with the sudden turns. Now let's talk about why, and what the paper suggests we do about it.

Jane: Right. So the researchers found a really interesting pattern. The typhoons that the AI model predicted well were the ones that were basically being pushed along by a strong, steady wind, the western Pacific subtropical high. It's like a conveyor belt.

Tom: And when that conveyor belt is strong, the typhoon just follows it. The AI model is great at predicting the big picture, the large-scale weather patterns. So it nails those cases.

Jane: But when the conveyor belt is weak, the typhoon's path depends much more on its own internal structure. The fine details of the storm itself. And that's where the AI model falls apart.

Tom: And the paper argues that's because AI models are trained to match historical data, but they don't actually simulate the physics. They don't have a real understanding of the storm's core, the updrafts, the downdrafts, the way the pressure and wind interact.

Jane: And they found evidence of that. The AI model actually failed to predict the strong vertical winds in the typhoon. It just smoothed them out. And it also had trouble with the relationship between wind and pressure, which is a fundamental physical law.

Tom: So it's like the AI model is a brilliant artist who can paint a beautiful landscape, but if you ask it to paint a single leaf in a storm, it just smears it. It doesn't understand how the leaf actually moves.

Jane: That's a great analogy. So what do they suggest? How do we fix this?

Tom: They suggest two main things. First, we need higher-resolution training data. If the AI model has never seen the fine details of a typhoon's core, it can't learn to predict them. So we need better reanalysis datasets.

Jane: And the second thing?

Tom: The second thing is to add physical constraints to the AI model. Instead of just letting it learn from data, we need to build in the laws of physics. So the model is forced to produce results that are physically possible.

Jane: So it's not just about throwing more data at it. It's about teaching it the rules of the game.

Tom: Exactly. And that's a really important insight, because it suggests that the future of weather forecasting isn't either AI or physics. It's a combination of both.

Jane: And that's a much more hopeful message than "AI is bad." It's saying AI has a specific weakness, and we know how to fix it.

Tom: Right. So let's wrap this up in our final segment.

Conclusion: Tom: So we've spent this whole episode on "AI Models Still Lag Behind Traditional Numerical Models in Predicting Sudden-Turning Typhoons." And I think the big takeaway is that AI weather models are incredibly powerful, but they have a specific blind spot.

Jane: And that blind spot is the rare, extreme events. The sudden turns, the unusual paths. The things that don't happen very often, so the AI model hasn't seen enough examples to learn them.

Tom: And the paper shows that for those cases, the traditional physics-based models are still the ones you want to trust. Especially when you need a forecast more than two or three days out.

Jane: But it's not a defeat. It's a roadmap. The paper gives us clear directions on how to improve AI models, with better data and physical constraints.

Tom: And that's the exciting part. This isn't the end of the story. It's a challenge. A challenge to the AI community to build models that don't just handle the average, but also the extreme.

Jane: And that's so important, because as the climate changes, we're likely to see more of these weird, unpredictable weather events. We need tools that can handle them.

Tom: So we're saying goodbye to this paper, but we're taking its message with us. AI is the future, but it's not the present. Not yet.

Jane: And that's a good thing to remember. We need to keep investing in the traditional models while we push the AI models to get better.

Tom: Alright, that's it for this one. Thanks for listening, everyone. We'll be back soon with another paper to dig into. See you then.

Jane: Bye everyone!

Daosheng Xu, Zebin Lu, Jeremy Cheuk-Hin Leung, Dingchi Zhao, Yi Li, Yang Shi, Bin Chen, Gaozhen Nie, Naigeng Wu, Xiangjun Tian, Yi Yang, Shaoqing Zhang, Banglin Zhang

Guangzhou Institute of Tropical and Marine Meteorology · Guangdong Provincial Key Laboratory of Regional Numerical Weather Prediction · China Meteorological Administration · Key Laboratory of Physical Oceanography · Ministry of Education · Institute for Advanced Ocean Study · Frontiers Science Center for Deep Ocean Multispheres and Earth System · College of Oceanic and Atmospheric Sciences · Ocean University of China · College of Meteorology and Oceanography · National University of Defense Technology · Guangdong Meteorological Observatory · College of Atmospheric Science · Lanzhou University · National Meteorological Centre · State Key Laboratory of Tibetan Plateau Earth System · Resources and Environment · Institute of Tibetan Plateau Research · Chinese Academy of Sciences

physics.ao-ph, cs.LG

Submitted: 2025-02-22

Updated: 2026-08-17

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 49/100

Key concepts

Sudden-Turning Typhoons
These are typhoons that make a sharp turn in their path, which is dangerous because they catch people off guard. The paper focuses on these extreme events as a specific area where AI models struggle to predict accurately compared to traditional models.
Pangu-Weather and ECMWF-IFS
Pangu-Weather is an AI model that was tested against the traditional European model, ECMWF-IFS. The study found that Pangu-Weather was more accurate for ordinary typhoons, while the traditional physics-based model outperformed it for sudden-turning typhoons.
Physical Constraints
This suggests adding the laws of physics into AI models. Instead of just learning from historical data, forcing the model to produce results that are physically possible helps it understand how storms move and interact with each other, addressing its weakness in simulating fine storm details.

Terminology

Summary

Summary

This paper evaluates the performance of the AI-based weather prediction (AIWP) model Pangu-Weather against traditional numerical weather prediction (NWP) models and human forecasters in predicting tropical cyclone (TC) tracks, with a specific focus on extreme and unusual TC trajectories, particularly sudden-turning typhoons. The authors state: "Given the interpretability, accuracy, and stability of numerical weather prediction (NWP) models, current operational weather forecasting relies heavily on the NWP approach. In the past two years, the rapid development of Artificial Intelligence (AI) has provided an alternative solution for medium-range (1 10 days) weather forecasting. Bi et al. introduced the first AI-based weather prediction (AIWP) model in China, named Pangu-Weather, which offers fast prediction without compromising accuracy. In their work, Bi23 made notable claims regarding its effectiveness in extreme weather predictions. However, this claim lacks persuasiveness because the extreme nature of the two tropical cyclones (TCs) examples presented in Bi23, namely Typhoon Kong-rey and Typhoon Yutu, stems primarily from their intensities rather than their moving paths. Their claim may mislead into another meaning which is that Pangu-Weather works well in predicting unusual typhoon paths, which was not explicitly analyzed."

The study reassesses Pangu-Weather's ability to predict extreme TC trajectories from 2020–2024. The authors divided 104 TC cases into three categories: ordinary TCs (83 cases), sudden-turning TCs (14 cases), and looping TCs (7 cases). Four types of prediction approaches were compared: (1) Pangu-Weather (including Pangu-ERA5, Pangu-ECMWF, and Pangu-NCEP); (2) global NWP models (ECMWF-IFS and NCEP-GFS); (3) a regional NWP model (CMA-TRAMS-L125); and (4) human forecasters (official forecasts from JTWC, JMA, and CMA).

The overall results show that "Pangu-Weather performs exceptionally well in predicting TCs’ trajectories, especially in 2023. Pangu-ERA5 exhibited greater accuracy than ECMWF-IFS, which is consistent with Bi23’s conclusion. Although the accuracies slightly drop when Pangu-Weather is driven by real-time ECMWF-IFS analysis, Pangu-ECMWF still outperforms ECMWF-IFS for most TCs. However, the authors found that Despite the overall better performance, Pangu-Weather had larger errors in predicting uncommon TC trajectories, such as Severe Typhoon Khanun (202306). It made two sharp turns within 5 days when it passed through the Ryukyu Islands and caused significant damage in surrounding countries. Despite its weaker intensity compared to Kong-rey and Yutu, Khanun’s exceptional path rendered it an extreme event, posing challenges for accurately predicting its movements and landfall position. In this case, Pangu-ERA5 and Pangu-ECMWF exhibit average 24–120 h track forecast errors 7.0% and 18.8% greater than those of ECMWF-IFS, respectively."

For Typhoon Khanun specifically, the paper reports: "Pangu-Weather demonstrated smaller average biases and uncertainties compared to human forecasters, but greater than NWP models, for the 24-to-120-hour track forecasts of Khanun. Furthermore, compared with both human forecasters and NWP models, Pangu-Weather exhibited a steeper decline in prediction skill as forecast lead time increased. Consequently, the difference between Pangu-ECMWF and human forecasters becomes indistinguishable for 5-day forecasts. These findings indicate Pangu-Weather’s limitations in predicting rarely-occurring extreme TC cases. Although Pangu-Weather demonstrates good prediction skills in global-scale forecasts, it only catches the performance of NWP models for short-term forecasts of 1-2 days in extreme cases like Typhoon Khanun, but not for longer lead times."

The paper extends the analysis to all 14 sudden-turning TCs: "We find that ECMWF-IFS overall outperforms Pangu-ECMWF for these sudden-turning cases. The mean distance error of Pangu-ECMWF ranges from 40.9–400.2 km for 6–120 forecast hours, which are on average 9.4% greater than those of ECMWF-IFS (ranging from 30.1–354.0 km). This further confirms the conclusion that Pangu-Weather performs worse, compared with NWP models, in capturing the extreme nature of these unusual TC paths. Conversely, for ordinary TCs, Pangu-ECMWF overall outperforms ECMWF-IFS... The mean distance error of Pangu-ECMWF (ranging from 49.9–244.9 km) is on average 13.5% smaller than those of ECMWF-IFS (ranging from 48.8–277.2 km). For looping TCs, Pangu-ECMWF also performs well... The mean distance error of Pangu-ECMWF (ranging from 52.8–240.5 km) is on average 32.9% smaller than those of ECMWF-IFS (ranging from 40.8–403.2 km). The authors note that Analyses based on other two AIWP models, Fengwu and FuXi, yield consistent results."

The paper investigates the key factor limiting AIWP models' forecast skills for sudden-turning TCs. The authors find that "TCs that are better predicted by ECMWF-IFS predominantly exhibit usual trajectories, following the large-scale climatological steering flow of the western Pacific subtropical high (WPSH). They are associated with background southeasterly winds and a stronger WPSH. Under the influence of a stronger WPSH, TCs’ moving trajectories are more likely to follow the background steering flow. In contrast, for TCs that are better predicted by Pangu-ECMWF, the guiding airflow is mainly southwesterly winds, and the WPSH is weaker. Under the background of a weaker WPSH, TCs’ trajectories are often the result of the interaction between the background steering flow and the TC itself, and result in anomalous paths. In this case, the prediction accuracies of the fine structural characteristics of the TC core region become important. The authors conclude that Since AIWP models, such as Pangu-Weather, are more capable of capturing the evolution patterns of large-scale circulations (e.g., WPSH), they tend to give better prediction results for TCs when the WPSH is strong. On the contrary, AIWP models are not good at 'simulating' TC structures and fail to reproduce the realistic physics of TC systems. Thus, AIWP models tend to give worse prediction results for the fine structural characteristics of the TC core region when the WPSH is relatively weak."

The paper discusses potential downsides of the data-driven approach: "Unlike NWP models which are designed to solve a set of partial differential equations governing the atmosphere, AIWP models are data-driven models trained to make predictions by learning historical weather patterns. This notably different approach may lead to several potential downsides limiting AIWP models’ ability to predict extreme events. For example, the fully data-driven training approach makes AIWP models more adept at capturing typical weather patterns, but less effective at predicting extreme cases. The limited samples of extreme events make AIWP models hard to predict patterns that rarely occurred in the past. Additionally, the authors note that Pangu-Weather exhibits apparent error in predicting the geostrophic relationship between geopotential and wind, as well as Khanun’s strong updraft and downdraft, although it well captures the TC warm core structure and its quasi-hydrostatic approximation. This evidence suggests that AIWP models have the potential to generate predictions that may be unrealistic and not inherently bound by physical laws. Besides, the use of global root-mean-square error as the model cost function tends to penalize predictions of extreme values, hence generating relatively smooth forecast outputs, which is also evident in the underestimation of TC intensities. In addition, the above issue could also smooth out atmospheric systems surrounding a TC and subsequently influence the prediction of TC moving speed and direction."

The authors conclude: "Our analyses indicate that a major issue affecting AIWP models’ prediction skills on extreme TC trajectories is their limited ability to forecast the fine structural characteristics of TCs, which lowers the ability of AIWP models to reasonably describe the interaction between TCs and large-scale steering flow when the WPSH is weak, thereby affecting the accuracy of sudden-turn predictions in TC trajectories. They propose two directions for improvement: (1) creating higher-resolution reanalysis datasets and training high-resolution AI models based on these datasets; (2) recognizing that traditional NWP models have advantages in forecasting extreme events, such as sudden turns in TC trajectories. Therefore, it is necessary to integrate more physical constraints into AIWP models, to mitigate the problem of excessive smoothing in the prediction of certain extreme events, a limitation often encountered in pure AI models with restricted sample sizes."

The final conclusion states: "While Pangu-Weather exhibits higher overall prediction skills in TC track predictions, it still lags behind ECMWF-IFS in predicting sudden-turning TCs, compared to ECMWF-IFS. Especially for extreme cases like Khanun, Pangu-Weather achieves high accuracy in 1-to-2-day forecasts, but not for longer lead times. This highlights the necessity of improving AIWP models’ medium-range forecasting skills under rare weather conditions. The AI-based Pangu-Weather model has already demonstrated advantages over human forecasters, demonstrating the power of AIWP models in extracting natural patterns from historical data. Moving forward, unraveling small sample size problems and incorporating physical constraints into the model training process could enhance AIWP models’ ability to predict abnormal weather phenomena, to cope with the increasing risks of extreme weather caused by climate change."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI weather prediction systems:

  • Current limitation: AI models (Pangu-Weather) fail to reproduce realistic TC core structures, geostrophic relationships, and vertical velocities

  • Improvement: Add physics-based loss terms during training that penalize violations of:

  • Geostrophic balance between geopotential height and wind fields

  • Quasi-hydrostatic relationships

  • Continuity equation constraints for vertical velocity

  • Result: Models will produce physically consistent outputs even for rare extreme events

  • Current limitation: Global RMSE cost function penalizes extreme values, producing smooth outputs that miss sudden-turning TC trajectories

  • Improvement:

  • Use weighted loss functions that amplify errors in extreme weather regions (TC cores, sudden direction changes)

  • Apply higher weights to samples with anomalous trajectories during training

  • Implement focal-loss-style weighting to focus on rare, high-impact events

  • Result: Better prediction of sudden-turning TCs (Khanun-type events) without sacrificing overall accuracy

  • Current limitation: AI models excel at ordinary TCs (13.5% better than ECMWF-IFS) but fail at sudden-turning TCs (9.4% worse)

  • Improvement:

  • Create ensemble system that combines AI predictions with NWP outputs

  • Use AI model when WPSH is strong (southeasterly steering flow)

  • Switch to NWP model when WPSH is weak (southwesterly flow) or when sudden-turning patterns are detected

  • Result: Optimal performance across all TC categories by leveraging strengths of both approaches

  • Current limitation: AI models trained on coarse data cannot resolve fine TC core structures

  • Improvement:

  • Train on higher-resolution reanalysis data (0.1° or finer)

  • Incorporate TC-specific datasets with detailed inner-core observations

  • Use transfer learning from regional high-resolution NWP outputs

  • Result: Better representation of TC dynamics that influence sudden directional changes

  • Current limitation: AI model skill degrades faster than NWP with increasing lead time (steep decline beyond 48h for extreme cases)

  • Improvement:

  • Train model with lead-time-dependent loss weighting

  • Add temporal consistency constraints across forecast steps

  • Implement ensemble forecasting with multiple initial condition perturbations

  • Result: Maintain prediction skill for 3-5 day forecasts of extreme TC trajectories

The enhanced AI system will:

  • Predict sudden-turning typhoons with accuracy comparable to ECMWF-IFS (reducing 24-120h track errors by 20% for Khanun-type events)

  • Maintain skill for 5-day forecasts of extreme events instead of degrading to human-forecaster levels

  • Generate physically consistent outputs including realistic vertical velocities, geostrophic winds, and TC warm core structures

  • Automatically switch strategies based on detected WPSH strength and TC trajectory patterns

  • Provide reliable predictions for all TC categories (ordinary, sudden-turning, and looping) with overall accuracy exceeding both pure AI and pure NWP approaches

  • Reduce operational forecasting errors for rare but destructive weather events, potentially saving millions in disaster mitigation costs

Abstract

Given the interpretability, accuracy, and stability of numerical weather prediction (NWP) models, current operational weather forecasting relies heavily on the NWP approach. In the past two years, the rapid development of Artificial Intelligence (AI) has provided an alternative solution for medium-range (1-10 days) weather forecasting. Bi et al. (2023) (hereafter Bi23) introduced the first AI-based weather prediction (AIWP) model in China, named Pangu-Weather, which offers fast prediction without compromising accuracy. In their work, Bi23 made notable claims regarding its effectiveness in extreme weather predictions. However, this claim lacks persuasiveness because the extreme nature of the two tropical cyclones (TCs) examples presented in Bi23, namely Typhoon Kong-rey and Typhoon Yutu, stems primarily from their intensities rather than their moving paths. Their claim may mislead into another meaning which is that Pangu-Weather works well in predicting unusual typhoon paths, which was not explicitly analyzed. Here, we reassess Pangu-Weather's ability to predict extreme TC trajectories from 2020-2024. Results reveal that while Pangu-Weather overall outperforms NWP models in predicting tropical cyclone (TC) tracks, it falls short in accurately predicting the rarely observed sudden-turning tracks, such as Typhoon Khanun in 2023. We argue that current AIWP models still lag behind traditional NWP models in predicting such rare extreme events in medium-range forecasts.

Sources

Related papers