Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods
Matthew Kahn, Milan Arjel, Nirmala Adhikari, Mingmar Sherpa, James Pope
Tribhuvan University · Cornell University · University of Alabama at Birmingham · University of Bristol
cs.LG, stat.AP
Submitted: 2026-08-12
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper tests whether free satellite data can predict glacial lake outburst floods (GLOFs), rainfall-triggered landslides, and smaller glacial floods in High Mountain Asia, with a focus on Nepal.
Terminology
Summary
This paper tests whether free satellite data can predict glacial lake outburst floods (GLOFs), rainfall-triggered landslides, and smaller glacial floods in High Mountain Asia, with a focus on Nepal. The study separates two questions: which site is susceptible
(based on static terrain features) and when a trigger arrives
(based on antecedent weather). Using 589 dated outbursts from HMAGLOFDB and thousands of catalogued landslides, the authors match each event against similar but unfailed sites and validate all models under spatial cross-validation that withholds whole map tiles.
The key results are:
-
Triggering (WHEN): Antecedent weather times the trigger at ROC 0.73 for big bursts, 0.83 for landslides, and 0.82 for small floods. The ordering is physically sensible:
the small floods and the landslides are the most directly rainfall- and melt-driven, while the big moraine bursts include the avalanche-triggered failures that no weather signal can see.
-
Susceptibility (WHICH-site): Terrain ranks susceptibility only in part. Scored naively across all of High Mountain Asia, terrain appears near-perfect (ROC 0.92 for big bursts, 0.91 for small floods, 0.98 for landslides), but this is a geographic artifact:
The catalogued events cluster in the wet, monsoon-fed ranges where floods are both frequent and observed, while the comparison pool spans the whole region, including the cold, dry Tibetan interior.
When matched against comparable nearby sites within 50 km, the honest figures are 0.76 for big bursts, 0.71 for landslides, and 0.54 (no better than chance) for small floods. -
Within-region signal: The burst signal holds within single regions, reaching 0.89 in Nepal alone, then 0.75 in Bhutan and the eastern Himalaya, 0.70 in the Karakoram, and 0.69 in the western Himalaya. The signal is mechanism-specific: ice-dammed lakes separate at 0.93, moraine-dammed bursts at 0.72, and supraglacial small floods not at all (0.58).
-
Deep learning comparison: Five deep-learning models (self-supervised pretraining, anomaly detection, attention, physics-informed, positive-unlabeled) do not decisively beat a simple gradient-boosted baseline. On landslides, three deep models score marginally higher (attention at 0.77, physics-informed at 0.76, self-supervised at 0.75 vs. baseline 0.71), but
the baseline's interval on landslides runs from 0.65 to 0.77, and all three leading models fall inside it.
For the lake hazards, the baseline wins outright. -
Interpretability: The winning burst model reduces to a three-rule decision tree on ruggedness, annual precipitation, and elevation, which reproduces the full model at 0.77 against the full model's 0.76. The Shapley attribution ranks
surrounding ruggedness first, then annual precipitation, then elevation.
Ablation confirms the ranking:Removing the top three features and refitting costs a tenth of a point of skill on the pooled bursts, and nearly two tenths within Nepal.
-
Deformation: The interferometric deformation signal
did not separate
burst lakes from dangerous-but-unfailed lakes:burst lakes and flagged-but-stable lakes deform alike, and a within-lake test found no acceleration before a lake's own failure beyond its own quiet periods.
Deformation isa necessary-not-sufficient hazard marker, to be fused with the weather trigger rather than used alone to rank lakes.
-
Deliverable: The paper closes with a ranked Nepal watchlist built from the within-Nepal burst model (ROC 0.89) run over 47 named lakes in ICIMOD's inventory of potentially-dangerous glacial lakes, overlaid with monsoon exposure. The list is
a prioritisation aid, not a prediction.
The paper's six contributions are: (1) a matched control that shows unmatched susceptibility scores are inflated (0.92 drops to 0.76); (2) a working separation of susceptibility from triggering; (3) a three-hazard benchmark under identical protocol; (4) an auditable three-rule susceptibility model; (5) a ranked, named Nepal watchlist; and (6) honest nulls for small floods, deformation ranking, and deep learning. The authors conclude that a free-data screening layer can honestly be built
on triggering for daily signals and susceptibility as a slow prior, with the caveat that the scarce quantity here is labeled failures and clean features, not model capacity.
Improvements for AI systems
Improvements to AI systems:
-
Add a spatial cross-validation protocol that withholds entire geographic tiles (e.g., 50 km blocks) during training. This prevents the AI from learning spurious correlations between event density and terrain features (e.g., mistaking
wet monsoon regions
forhigh hazard
). The improved system will report honest, generalizable performance metrics instead of inflated ones (ROC drops from 0.92 to 0.76 when this is applied). -
Separate the hazard model into two distinct modules: a static susceptibility module (terrain, elevation, ruggedness) and a dynamic triggering module (antecedent weather, melt, rainfall). The AI will output two scores per site—
where
andwhen
—instead of a single risk score. This allows users to prioritize monitoring during trigger windows and to rank sites for long-term planning independently. -
Implement a matched-control sampling strategy for training data. Instead of comparing failed sites against all non-failed sites across a vast region, the AI will pair each failure with similar nearby sites (within 50 km, matched on elevation, slope, and lake type). This eliminates geographic bias and produces honest susceptibility rankings, especially for rare events like small glacial floods (which drop from 0.91 to 0.54—chance—when matched properly).
-
Use a three-rule decision tree as a default interpretable baseline for susceptibility (ruggedness > annual precipitation > elevation). The improved system will always train this simple model first, and only escalate to complex models if they beat it by a statistically significant margin. This ensures the AI remains auditable and deployable in low-resource settings, while still matching the performance of black-box models (0.77 vs. 0.76 ROC).
-
Integrate a physics-informed trigger prior for weather-driven hazards. The AI will weight antecedent precipitation and melt more heavily for landslide and small-flood predictions (ROC 0.83 and 0.82), but down-weight weather for moraine-dammed lake bursts (ROC 0.73) because those include avalanche triggers invisible to weather data. This prevents the system from over-relying on weather when it cannot explain the failure mode.
-
Add a
necessary-not-sufficient
fusion layer for deformation data. Instead of using InSAR deformation as a standalone ranking feature (which failed to separate burst lakes from stable ones), the AI will use deformation only as a gating condition—if deformation is absent, the site is deprioritized; if present, the AI then applies the weather trigger and terrain susceptibility to estimate imminent risk. This prevents false alarms from deformation alone. -
Build a region-aware calibration layer that adjusts susceptibility scores based on geographic sub-region (Nepal, Bhutan, Karakoram, etc.). The improved system will report separate confidence intervals per region (e.g., Nepal ROC 0.89 vs. western Himalaya 0.69) and will not extrapolate a model trained in one region to another without retraining or domain adaptation.
-
Implement a
null-result reporting
feature in the model evaluation pipeline. The AI will automatically flag when a model's performance is not statistically better than chance (e.g., small floods at 0.54) or when deep learning fails to beat a gradient-boosted baseline. This prevents overclaiming and forces the system to outputno reliable signal
rather than a confident but meaningless prediction. -
Create a prioritized watchlist generator that combines the within-region susceptibility model with a monsoon exposure overlay. The improved system will output a ranked list of named lakes (e.g., 47 from ICIMOD) with a clear label:
prioritization aid, not prediction.
Each entry will include the three-rule explanation (ruggedness, precipitation, elevation) and the current trigger score from the weather module, enabling field teams to allocate resources efficiently. -
Add a data-scarcity-aware training mode. The AI will explicitly quantify the uncertainty from limited labeled failures (e.g., 589 outbursts) and will avoid overfitting by using simple models (gradient boosting, decision trees) when data is sparse. It will also flag when adding more labeled failures or cleaner features would yield more improvement than increasing model complexity—directly addressing the paper's conclusion that
the scarce quantity is labeled failures, not model capacity.
Abstract
Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slowly sagging, and satellite weather marks the weeks when a primed lake is under stress. A companion feasibility study found that deformation indicates which lake is destabilizing and weather indicates when it is at risk, but proposed no predictive model. To address this gap, we propose and evaluate models that predict which site is susceptible and when a trigger arrives. We test three related hazards on free data alone: large moraine- and ice-dammed bursts, rainfall-triggered landslides, and smaller floods from ponds on and around a glacier. Each hazard gets two questions, never blended. Using 589 dated outbursts from HMAGLOFDB and several thousand catalogued landslides, we match each event against similar but unfailed sites, and hold every model to a strong simple baseline under spatial cross-validation that withholds whole map tiles, so no model succeeds by recognising a trained-on neighbourhood. Antecedent weather times the trigger at ROC 0.73 for big bursts, 0.83 for landslides, and 0.82 for small floods. Terrain ranks susceptibility only in part: scored naively it appears near 0.9, largely because catalogued failures cluster in wetter ranges; matched against comparable nearby sites the honest figures are 0.76, 0.71, and 0.54 (no better than chance). The burst signal holds within single regions, reaching 0.89 in Nepal alone. Five deep-learning models do not decisively beat a simple gradient-boosted baseline. Three score marginally higher on landslides, a hint too small to confirm. For the lake hazards the baseline wins outright, reproduced by a three-rule decision tree on ruggedness and monsoon rainfall. We close with a ranked Nepal watchlist, a prioritisation aid, not a prediction, and note where free data reaches its limits.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks