Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval

arXiv:2510.20486 · cs.LG, cs.AI, physics.ao-ph, physics.geo-ph · Submitted 2025-10-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval".

Jane: The paper was written by Fangjian Zhang, Xiaoyong Zhuge, Wenlan Wang, Haixia Xiao, Yuying Zhu et al. from Nanjing Innovation Institute for Atmospheric Sciences and Chinese Academy of Meteorological Sciences and Jiangsu Meteorological Service and Jiangsu Key Laboratory of Severe Storm Disaster Risk and Key Laboratory of Transportation Meteorology of China Meteorological Administration and Institute of Tibetan Plateau Meteorology and Heavy Rain and Drought-Flood Disasters in Plateau and Basin Key Laboratory of Sichuan Province and China Meteorological Administration.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are checking out a fascinating new paper today called "Hurdle–IMDL: An Imbalanced Learning Framework for Infrared Rainfall Retrieval."

Jane: That title definitely sounds technical, Tom, but it's really about helping AI get better at predicting heavy rain from satellites.

Tom: It sounds like the authors are tackling a massive gap in how we currently monitor the weather.

Lu: They are, because most AI models today struggle when the data they see doesn't look like the rare, intense storms that actually cause damage.

Jane: You're talking about that imbalance issue mentioned in the title, right, Lu?

Lu: Exactly, because the data is mostly just "no rain" or "light rain," which makes the heavy stuff look like an outlier to the computer.

Meng: I wonder if the team behind this has already dealt with the messy reality of real-world sensor data.

Tom: They definitely have, since the authors, like Fangjian Zhang and Xiaoyong Zhuge, are coming from the Nanjing Innovation Institute for Atmospheric Sciences.

Jane: It's great to see researchers from the Chinese Academy of Meteorological Sciences working on this, because they have the actual ground truth data.

Meng: Having that connection to actual meteorological services must make the implementation much more practical for real weather stations.

Lalam: This kind of work is vital because if we can't accurately predict extreme weather, our entire global approach to disaster preparedness stays stuck in the past.

Tom: It really does feel like they're building a bridge between pure AI theory and actual atmospheric physics.

Jane: And that bridge is exactly what they use to fix the way we look at infrared satellite signals.

Tom: Let's look closer at how they actually structure this "Hurdle" approach.

Summary: Jane: The researchers use what they call a "divide-and-conquer" strategy to handle the messy rain data.

Tom: Right, they basically split the problem into two separate challenges: the "zero inflation" and the "long tail."

Jane: Can you explain what those terms actually mean for someone who isn't a statistician, Tom?

Tom: Sure, zero inflation just means there are a ton of moments where it isn't raining at all, so the model gets overwhelmed by zeros.

Jane: And the long tail part is when you have plenty of light drizzle but very few examples of massive downpours.

Lu: That's where the "Hurdle" part comes in, where they use one part of the model to clear that hurdle of zero versus rain.

Tom: Then they use this new IMDL method to handle the actual amount of rain once they know it's happening.

Lu: The math behind the IMDL is actually quite beautiful because it uses Bayes' theorem to find an "ideal" model.

Meng: I'm curious about how they actually implement that without the model just getting lost in the complexity.

Lu: They assume the underlying physical process is constant, which lets them transform a biased model into an unbiased one.

Meng: So they aren't just guessing, they're using the physics of how rain forms to guide the AI?

Jane: It seems like they're using a lognormal distribution to model how the rain rates actually behave.

Lalam: By incorporating these physical constraints, they are teaching the AI to respect the laws of nature rather than just chasing patterns in a lopsided dataset.

Tom: It's a much smarter way to train than just throwing more data at a standard neural network.

Jane: It really changes the game for how we handle these imbalanced environmental variables.

Tom: Let's see if these theoretical improvements actually show up in the real-world testing.

Improvements: Tom: The results from their testing against other models like Diffusion and MTCF are pretty striking.

Jane: They specifically focused on how much the models underestimate heavy rain, which is a huge problem for flood warnings.

Tom: Most of the older models had a negative Mean Error, meaning they were consistently predicting less rain than what actually fell.

Jane: But the Hurdle–IMDL framework managed to bring that error way down, especially for those extreme events.

Meng: I saw in the paper that they tested it against several thresholds, even up to thirty millimeters per hour.

Tom: Yeah, and for those extreme thirty mm thresholds, the other models basically had zero predictive skill.

Jane: While the Hurdle–IMDL kept performing well, which is a massive jump in reliability.

Lu: I was looking at the case studies they included, like that Meiyu front event in July two thousand twenty-one.

Tom: That was a great example, because the other models just couldn't capture the intensity of the rainband.

Lu: They really struggled with the spatial extent, but the Hurdle–IMDL actually mapped out the heavy rain areas quite accurately.

Meng: It's impressive that it also handled the convective cells, where the rain is more scattered and harder to catch.

Jane: It seems like the ability to reduce that systematic underestimation is the biggest win here.

Lalam: This means we can move toward a culture of much more proactive emergency response because our data is finally catching the extremes.

Tom: It's a huge step forward for anyone relying on satellite data for safety.

Jane: We've seen some incredible breakthroughs here today.

Conclusion: Tom: We've spent a lot of time on "Hurdle–IMDL: An Imbalanced Learning Framework for Infrared Rainfall Retrieval," and it's clear this is a big deal.

Jane: It's a clever way to use physics and math to solve a very frustrating data problem in meteorology.

Lu: I think the next step is seeing how this can be applied to other variables like snow depth or even air pollution.

Meng: From my side, I'm looking forward to seeing how engineers can integrate this into real-time satellite processing pipelines.

Lalam: Ultimately, this research helps us build a more resilient society by making the invisible, extreme events visible to our technology.

Tom: That's a perfect place to stop. Thanks for joining us, everyone.

Jane: See you next time!

Nanjing Innovation Institute for Atmospheric Sciences · Chinese Academy of Meteorological Sciences · Jiangsu Meteorological Service · Jiangsu Key Laboratory of Severe Storm Disaster Risk · Key Laboratory of Transportation Meteorology of China Meteorological Administration · Institute of Tibetan Plateau Meteorology · Heavy Rain and Drought-Flood Disasters in Plateau and Basin Key Laboratory of Sichuan Province · China Meteorological Administration

cs.LG, cs.AI, physics.ao-ph, physics.geo-ph

Submitted: 2025-10-23

Updated: 2026-09-11

Comments: 19 pages

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: This paper presents the Hurdle–Inversion Model Debiasing Learning (IMDL) framework, designed to overcome the challenges of imbalanced label distribution in infrared rainfall retrieval.

Key concepts

Zero Inflation
This describes a statistical issue where a dataset contains an overwhelming number of zero values, such as moments when no rain is falling. For AI models, this abundance of 'no rain' data can make it difficult to accurately identify and predict when actual precipitation begins.
Long-Tailed Imbalance
This refers to a situation where common events, like light drizzle, dominate the data, while rare but critical events, like heavy downpours, are infrequent. Because extreme storms appear as outliers in the dataset, standard AI models often struggle to predict their intensity accurately.
Hurdle-IMDL Framework
This is a two-part strategy that separates rainfall prediction into two challenges. It first uses a 'hurdle' model to distinguish between no rain and rain, then applies an imbalanced learning method using Bayes' theorem and physical constraints to accurately predict the intensity of the rainfall.

Terminology

Summary

This paper presents the Hurdle–Inversion Model Debiasing Learning (IMDL) framework, designed to overcome the challenges of imbalanced label distribution in infrared rainfall retrieval. Because rainfall is characterized by extreme imbalances, conventional AI models often suffer from underfitting of AI models for rare events, which necessitates a new approach to accurately retrieve high-impact, heavy-to-extreme precipitation.

The dual nature of imbalance

The study adopts a divide-and-conquer strategy to address the complexities of rainfall distribution. The researchers argue that the inherent imbalance in environmental variables can be decomposed into two primary components:

  • Zero inflation, which is characterized by the predominance of non-rain samples.

  • Long tail, defined by the disproportionate abundance of light-rain samples relative to heavy-rain samples.

By separating these issues, the framework can apply specific mathematical treatments to both the occurrence of rain and its subsequent intensity, rather than attempting to solve a single complex problem.

The Hurdle–IMDL mechanism

To manage zero inflation, the framework utilizes a statistically robust hurdle model that separates occurrence probability modeling from rain rate modeling. To address the long-tail issue, the authors propose IMDL, which aims to fundamentally resolve the long-standing long-tail challenge by transforming the learning object into an unbiased ideal inverse model.

This is achieved by leveraging a key invariance: the forward model is determined primarily by the underlying physical process and remains independent of data acquisition. By applying a probability transformation, IMDL links the biased inversion model found in long-tailed datasets to an ideal version. This allows the model to learn directly from naturally collected data while mitigating the bias typically introduced by imbalanced distributions.

Mathematical optimization

The framework models the conditional distribution of rain rates using a lognormal distribution. The training process is driven by a negative log-likelihood (NLL) objective function that admits an analytical solution. This objective is decomposed into four distinct components to facilitate efficient learning:

  1. A Dry term (DryT i) for non-rain samples where Ri = 0.

  2. A Wet term (WetT i) for rain occurrence where Ri > 0.

  3. A LogNorm term (LogNormT i) to fit the rain rate distribution.

  4. A CorrT correction term that embodies IMDL’s adjustment and redirection of conventional learning.

Experimental validation

Comprehensive evaluations using Himawari-8 satellite data confirm that Hurdle–IMDL is superior to conventional, cost-sensitive, generative, and multi-task learning methods. The framework's primary strength lies in its ability to mitigate systematic underestimation and improve the retrieval of rare events.

The study highlights several key performance advantages:

  • A markedly lower RMSE for heavy rain samples compared to all baselines.

  • The highest Equitable Threat Scores (ETS) across all rain grades, particularly for heavy-to-extreme rain.

  • Improved spatial and intensity accuracy in case studies involving Meiyu fronts and convective cells, where other models like Diffusion or MTCF tended to overestimate rain areas or underestimate intensities.

Improvements for AI systems

1. Implementation of the Hurdle–IMDL Architecture for Skewed Environmental Regression

  • Improvement: Apply the divide-and-conquer strategy—decomposing imbalanced distributions into a zero-inflation component (handled by a continuous probabilistic hurdle model) and a long-tail component (handled by IMDL)—to all quantitative remote sensing tasks involving skewed variables, such as aerosol optical depth (AOD), trace gas concentrations, and sea surface temperature.

  • Capability: The system will transition from mean-biased predictions to event-accurate retrievals, specifically enabling the detection and precise quantification of rare, high-impact extreme events (e.g., pollution spikes or heatwaves) that conventional models typically underestimate.

2. Integration of the Analytical Negative Log-Likelihood (NLL) Objective Function

  • Improvement: Replace standard Mean Squared Error (MSE) or heuristic cost-sensitive weighting with the derived analytical NLL objective function that incorporates the ratio of empirical distribution to ideal distribution (F(R) over F(R')).

  • Capability: The system can learn an unbiased ideal inversion model directly from naturally collected, imbalanced datasets. This eliminates the need for computationally expensive generative modeling (like Diffusion models) or manual data resampling, while fundamentally correcting the systematic underestimation of extreme values.

3. Deployment of Single-Network Multi-Output Architectures for Joint Parameter Estimation

  • Improvement: Shift from two-stage/sequential modeling (detecting occurrence then estimating rate) to a single, end-to-end multi-output architecture (e.g., a modified U-Net) that jointly estimates the occurrence probability (p) and the distribution parameters (mu).

  • Capability: The system will leverage feature-sharing and mutual reinforcement between detection and estimation tasks. This reduces error propagation inherent in two-stage processes and improves spatial precision by allowing the detection branch to inform the intensity estimation branch via shared deep features.

4. Development of a Dynamic Shape Parameter (sigma) Estimation Module

  • Improvement: Advance the current framework by replacing the static hyperparameter selection of sigma with a dynamic estimation mechanism (e.g., an auxiliary neural network head) within the NLL framework to capture varying volatility across different intensity scales.

  • Capability: The system will achieve differentiated, non-uniform estimation of signal volatility, preventing the model from converging uniformly to upper bounds and allowing for more granular and accurate modeling of the varying uncertainty associated with extreme vs. light events.

Sources

Related papers