Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery

arXiv:2608.06406 · cs.CV, cs.LG · Submitted 2026-08-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery".

Jane: The paper was written by the authors from LMU Munich and TU Munich and Munich Center for Machine Learning.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone! I'm Tom, and as always, I'm here with my co-host, Jane. Today we're digging into a fresh arXiv paper called "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery."

Jane: And Tom, I have to say, this one caught my eye immediately. It's from a team at LMU Munich and TU Munich, including Laura Bader, Muhammad Ammar Ahmed, Xiao Xiang Zhu, and Göran Kauermann. They're tackling something really practical — measuring how tall trees are from space.

Tom: Measuring trees from space sounds almost too simple when you say it like that, Jane. But the twist here is that they're not just predicting tree height — they're also telling you how sure they are about that prediction. And that's a big deal for anyone making decisions based on this data.

Jane: Exactly. Think about it — if you're a government agency deciding where to invest in reforestation, or a company calculating carbon credits, you need to know whether a prediction is solid or just a guess. This paper is about giving you that confidence information alongside the actual numbers.

Tom: And the title mentions "sparse" forest height estimation. That's because they're working with tree inventory data, which is like having a few scattered ground measurements in a huge forest. Most pixels in their satellite images don't have a label at all. So they had to figure out how to train a model when most of the data is missing.

Jane: Right, and that's where the "deep evidential regression" part comes in. I'll keep it simple — it's a way for a neural network to output not just a single number, but a whole distribution of possible values. From that, you can calculate two types of uncertainty: one from noise in the data itself, and one from what the model doesn't know.

Tom: And that's what makes this paper exciting. They're applying a fairly new technique to a real-world problem with real satellite data — Sentinel-one and Sentinel-two — over the state of Bavaria in Germany. It's not just a toy example; it's a large-scale benchmark called TreeUQ.

Jane: So the big picture here is that we're moving from "here's a prediction" to "here's a prediction and here's how much you should trust it." That's a huge step for environmental monitoring, where mistakes have real consequences.

Tom: I love that framing, Jane. And in the next segment, we're going to get into the actual method they used and why it's so clever. Stick around!

Summary: Tom: Welcome back! We're still talking about "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery." Jane, let's break down what these researchers actually did.

Jane: So the core problem is this: they have satellite images of forests, and they want to predict the average tree height in each ten-meter pixel. But they only have actual measurements for a tiny fraction of those pixels — the ones where someone went out and measured trees on the ground.

Tom: And that's the sparse part. Most deep learning models would struggle with that because they need labeled data to learn from. But the authors came up with a clever fix — they introduced something called a masked evidential loss. Basically, when the model is training, it only looks at the pixels where it has a real label, and it ignores everything else.

Jane: That makes sense. But the more interesting part is what the model outputs. Instead of just giving a single height prediction, it outputs four parameters that define a whole probability distribution. From those, you can calculate the expected height, plus two types of uncertainty.

Tom: And those two types are aleatoric and epistemic uncertainty. Aleatoric is the randomness that's baked into the data — like sensor noise or the fact that trees in a pixel naturally vary in height. Epistemic is the model's own ignorance — what it doesn't know because it hasn't seen enough data from that region.

Jane: Exactly. And the model they used is a U-Net, which is a popular architecture for image-to-image tasks. They used a ResNet-fifty backbone that was pre-trained on ImageNet, and they fed it forty-nine channels of multimodal satellite data — that's a combination of radar data from Sentinel-one and optical data from Sentinel-two.

Tom: Now, the key result — and this is what got me excited — is that their uncertainty-aware model performed almost as well as a standard deterministic U-Net. The deterministic one got an RMSE of five point seven five meters, while the evidential one got five point nine one meters. That's a tiny difference, but the evidential model gives you all that extra uncertainty information for free.

Jane: And that's the trade-off that matters. You're not sacrificing much accuracy, but you're gaining a whole new dimension of insight. The paper also shows that the uncertainty estimates are well-calibrated — meaning when the model says it's ninety percent confident, it's actually right about ninety percent of the time.

Tom: So this isn't just a theoretical exercise. It's a practical method that works on real satellite data and gives you trustworthy uncertainty numbers. In the next segment, we'll talk about what improvements they suggest and where this could go next.

Improvements: Tom: We're back with "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery." Jane, I know you had some thoughts on what this paper improves compared to previous approaches.

Jane: Absolutely, Tom. The biggest improvement is efficiency. Traditional ways of getting uncertainty estimates — like training an entire ensemble of models or running the same model many times with random dropout — are computationally expensive. This approach gives you uncertainty in a single forward pass. One run through the network, and you get both the prediction and the uncertainty.

Tom: That's huge for real-world applications. If you're processing satellite imagery for an entire country, you can't afford to run ten models for every patch. This is a lightweight solution that scales.

Jane: And there's another improvement that I find really elegant. The paper introduces a masked loss function that handles the sparse labels gracefully. Instead of trying to fill in missing values or ignoring large parts of the image, the model simply focuses its learning on the pixels where it has ground truth. That's a clean solution to a messy real-world problem.

Tom: Let's bring in Lu, our senior AI researcher, to get a deeper perspective. Lu, what do you think is the most significant improvement here?

Lu: Thanks, Tom. I think the most significant improvement is how they validate the uncertainty estimates. They don't just say "here are some uncertainty numbers" — they actually check them against real forest structure data. The TreeUQ dataset includes per-pixel variance of tree heights from the inventory data. So they can ask: does the model's uncertainty go up when the forest is actually more heterogeneous?

Tom: And does it?

Lu: Yes, and quite strongly. The epistemic uncertainty — the model's lack of knowledge — correlates with the inventory-derived variance at zero point seven four Pearson correlation. That's a strong signal that the uncertainty isn't just noise; it's actually capturing something real about the forest structure.

Jane: That's the kind of validation that makes this paper stand out. They're not just reporting metrics on a test set; they're connecting the uncertainty to physical properties of the forest. That gives me confidence that this method could generalize to other regions.

Tom: And that's what we'll dig into next — the first page of the paper and what it sets up. Stay with us!

First Page: Tom: Welcome back to our discussion of "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery." Jane, let's go back to the very beginning — the abstract and introduction — because they set up the whole problem beautifully.

Jane: The abstract makes a strong case for why this matters. It talks about carbon accounting, biodiversity monitoring, and ecosystem management. These aren't abstract academic concerns — they're pressing global issues. And the paper argues that accurate forest height estimation is essential for all of them.

Tom: And then it drops this key phrase: "sparse supervision and geographic distribution shift." That's the real-world challenge. You train a model on one part of Bavaria, but forests in another part might look different. The model needs to know when it's out of its depth.

Jane: The introduction also references the TreeUQ benchmark, which was designed specifically for this problem. It provides multimodal satellite imagery — Sentinel-one radar and Sentinel-two optical — along with sparse tree inventory labels. And the benchmark explicitly motivates uncertainty-aware modeling, but the existing baselines were all deterministic.

Tom: So this paper is filling a gap that the benchmark creators themselves identified. They said "we need uncertainty," and these authors said "here's how you do it."

Jane: Let me bring in Meng, our lead engineer, because I think there's a practical angle here that's worth exploring. Meng, what do you see as the biggest practical challenge in implementing this?

Meng: Thanks, Jane. The biggest challenge I see is the computational cost of training. They used a ResNet-fifty encoder with forty-nine input channels, which is heavy. And they trained with eight random seeds to get reliable statistics. That's not cheap. But the inference is fast — one forward pass gives you everything. So for deployment, it's actually very practical.

Tom: That's a good point. Training is a one-time cost, but inference is what you do every day when you're processing new satellite imagery.

Meng: Exactly. And the fact that they used mixed-precision training and a cosine learning rate schedule shows they were thinking about efficiency from the start. This isn't a research prototype that can't run in production; it's designed to be usable.

Jane: And that brings us to the conclusion of our discussion. Let's wrap this up.

Conclusion: Tom: Alright, we're wrapping up our look at "Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery." Jane, what's your final takeaway?

Jane: My takeaway is that this paper shows you can have your cake and eat it too. You get accurate predictions — nearly as good as a deterministic model — plus you get well-calibrated uncertainty estimates, all in a single forward pass. That's a rare combination in deep learning.

Tom: And let's not forget the practical impact. This could help carbon credit markets work more honestly, help governments plan reforestation more effectively, and give ecologists better tools for monitoring biodiversity. When you know how confident a prediction is, you can make better decisions.

Jane: The authors also acknowledged limitations, which I appreciate. The Gaussian likelihood assumption might not perfectly capture forest height distributions, and the sparse supervision means the model might make spurious predictions in non-forest areas. But these are starting points for future work, not fatal flaws.

Tom: Let's give Lalam the final word. Lalam, what's the most impactful vision you see coming out of this?

Lalam: Thank you, Tom. The most impactful vision is a global forest monitoring system where every prediction comes with a confidence score that decision-makers actually trust. This paper is a step toward that. It shows that uncertainty quantification doesn't have to be an afterthought — it can be built into the model from the start. And as these methods mature, they could transform how we manage natural resources, how we verify climate commitments, and how we hold ourselves accountable for the health of our planet.

Tom: Beautifully said, Lalam. And with that, we're saying goodbye to this paper. Thanks to everyone who tuned in, and we'll see you next time with another exciting piece of research.

Jane: Until then, keep asking questions and stay curious. Goodbye, everyone!

LMU Munich · TU Munich · Munich Center for Machine Learning

cs.CV, cs.LG

Submitted: 2026-08-03

Updated: 2026-09-14

Code: https://github.com/ammarlam10/evidential

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 64/100

The gist: "Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management.

Key concepts

Deep Evidential Regression
A technique where a neural network outputs not just a single prediction but a distribution of possible values. This allows the model to quantify two types of uncertainty: noise in the data itself (aleatoric) and what the model does not know (epistemic).
Sparse Forest Height Estimation
The problem of predicting tree height using satellite imagery where ground measurements are only available for a small fraction of pixels. The method uses a masked evidential loss to train the model effectively despite this lack of labels.
Aleatoric Uncertainty
The randomness inherent in the data, such as sensor noise or natural variation in tree heights across different locations. This uncertainty is baked into the data and cannot be reduced by more training.
Epistemic Uncertainty
The model's own ignorance regarding a prediction, stemming from insufficient training data for a specific region. This uncertainty reflects what the model does not know because it hasn't seen enough relevant information.

Terminology

Summary

Summary

The paper investigates Deep Evidential Regression (DER) for uncertainty-aware forest height estimation from multimodal satellite imagery, evaluated on the TreeUQ benchmark. The authors state: "Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management. While recent deep learning approaches provide accurate predictions, they typically do not quantify predictive uncertainty. This limitation is particularly relevant in geospatial settings characterized by sparse supervision and geographic distribution shift."

The problem is framed as follows: "Estimating tree heights from satellite imagery is inherently uncertain due to sensor noise, heterogeneous vegetation structure, and geographic distribution shift. These challenges are amplified in sparse supervision settings, where only a small subset of pixels contains valid annotations. The TreeUQ benchmark provides multimodal satellite imagery together with sparse tree inventory labels for forest structure estimation at 10 m resolution for the region of Bavaria in Germany, and explicitly motivates uncertainty-aware modeling, yet existing baseline methods remain fully deterministic."

The contributions are threefold: "(1) We adapt DER to sparse multimodal Earth observation regression for forest height estimation. (2) We introduce a masked evidential loss formulation that enables DER model training under extreme geospatial label sparsity. (3) We provide a comprehensive empirical evaluation of predictive uncertainty, including calibration analysis, spatial uncertainty characterization, and an investigation of the relationship between uncertainty, forest structural heterogeneity, and prediction error."

Methodologically, the authors use a U-Net architecture with a ResNet-50 encoder pre-trained on ImageNet, adapted to 49 input channels from stacked Sentinel-1 SAR and Sentinel-2 optical composites. Input patches are 128×128 pixels, and the target is mean tree height per 10 m pixel. The deterministic U-Net is trained with a masked Smooth L1 loss, while the evidential U-Net replaces the scalar regression head with an evidential output layer of four neurons (zγ, zν, zα, zβ) that predicts Normal-Inverse-Gamma (NIG) hyperparameters m = (γ, ν, α, β). The evidential model predicts mean tree height as γ, with aleatoric uncertainty u2al = β/(α−1), epistemic uncertainty u2ep = β/(ν(α−1)), and total predictive uncertainty u2pred = u2al + u2ep. The authors introduce a masked evidential loss: "we define a binary patch-wise validity mask δ ∈ 0, 1 H×W, where δij = 1 indicates that a valid tree height label is available at spatial location (i, j) within the patch. Then, during training, loss computation is restricted to valid pixels only." The masked loss is computed as LmDER(w; y, m) = (1/Σδij) Σ δij · LDER(w; yij, mij), where LDER is the original evidential loss from Amini et al. [2].

Experimental setup follows the official TreeUQ benchmark with geographically disjoint train, validation, and test regions. A log(1+y) transformation is applied to stabilize optimization. Both models share identical preprocessing, masking, optimization, and evaluation settings. Each model is trained independently using eight random seeds with fixed hyperparameters: AdamW optimizer, learning rate 10−4, weight decay 10−4, cosine annealing schedule, mixed-precision training, batch size 32. Data augmentation includes random horizontal/vertical flips (p=0.5) and random 90° rotations (p=0.75), applied jointly to inputs, targets, and masks. Validation is performed every five epochs, with the best checkpoint (lowest validation RMSE) retained. Early stopping occurs if no improvement for ten consecutive validation steps, with a maximum of 100 epochs.

Results on the held-out test set show: the DER U-Net attains an RMSE of 5.91 ± 0.15m compared to 5.75 ± 0.09m for the deterministic baseline. The deterministic U-Net achieves MAE of 4.40 ± 0.06m and R2 of 0.525 ± 0.015, while the DER U-Net achieves MAE of 4.53 ± 0.15m and R2 of 0.497 ± 0.026. The authors note: the performance differences are small relative to the observed variability across random seeds, and incorporating evidential uncertainty estimation does not substantially degrade predictive performance while additionally providing uncertainty estimates in a single forward pass.

Calibration is assessed using prediction interval coverage following Kuleshov et al. [12]. The calibration curve closely follows the ideal diagonal, indicating that the predicted uncertainties are generally well calibrated, with a low Expected Calibration Error (ECE) of 0.025. The authors note: the model tends to be slightly overconfident for coverage levels below 70%, while exhibiting mild underconfidence for larger prediction intervals.

The paper also analyzes the relationship between predicted uncertainty and inventory-derived tree height variance statistics. At the individual-pixel level, no meaningful relationship is observed, but after aggregation to the patch level, clear positive associations are visible. Specifically: "the inventory-derived variance exhibits a moderate correlation with the predicted aleatoric uncertainty (Pearson r = 0.53, Spearman ρ = 0.55). Interestingly, an even stronger relationship is observed for epistemic uncertainty (r = 0.74, ρ = 0.76). Additionally, patch-wise prediction errors increase with the inventory-derived variance (r = 0.58, ρ = 0.60). The authors interpret this as: structurally heterogeneous forest stands are not only associated with intrinsic variability in the target variable, but also with increased model uncertainty."

Qualitative spatial analysis shows: Higher epistemic uncertainty is observed near transitions between forested and non-forested areas, whereas uncertainty is generally lower within homogeneous forest stands. The authors also highlight a limitation: "since the masked training loss is evaluated only on pixels with valid tree inventory annotations, the model is never explicitly trained to predict zero height in non-forest regions. Consequently, both the deterministic and evidential models occasionally produce spurious tree-height predictions in background areas."

Limitations acknowledged include: "While DER provides a decomposition into aleatoric and epistemic uncertainty, the interpretation of these quantities remains heuristic and should be treated with caution [16]. Moreover, the evidential framework relies on a Gaussian likelihood assumption, which may not fully capture the distribution of forest heights despite the applied log-transformation. Future work should compare DER with alternative uncertainty estimation approaches, and extend the analysis to additional forest structure variables and Earth observation benchmarks." The code is made available at https://github.com/ammarlam10/evidential.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:

  1. Add an evidential output head to the U-Net architecture (replacing the scalar regression head with 4 output neurons: γ, ν, α, β) to enable single-forward-pass uncertainty estimation.

  2. Implement the masked evidential loss function (Equation 1) that restricts loss computation to pixels with valid tree inventory annotations, preventing missing labels from corrupting uncertainty estimates.

  3. Add the NIG parameterization constraints: γ = zγ, ν = softplus(zν), α = softplus(zα) + 1, β = softplus(zβ) to ensure valid Normal-Inverse-Gamma distribution parameters.

  4. Compute aleatoric and epistemic uncertainty from the evidential parameters using: u2 al = β/(α−1) and u2 ep = β/(ν(α−1)).

  5. Add a calibration evaluation module that constructs prediction intervals from predictive mean and total uncertainty, then computes Expected Calibration Error (ECE) to verify uncertainty quality.

  • Predict forest height at 10m resolution from Sentinel-1/-2 imagery with accuracy comparable to deterministic models (RMSE ≈ 5.91m vs 5.75m baseline).

  • Output per-pixel uncertainty estimates (both aleatoric and epistemic) in a single forward pass, without needing ensembles or Monte Carlo dropout.

  • Provide well-calibrated prediction intervals (ECE = 0.025), meaning users can trust that a 90% confidence interval actually contains the true height 90% of the time.

  • Identify structurally heterogeneous forest regions where predictions are less reliable, since epistemic uncertainty correlates strongly (r = 0.74) with within-pixel tree-height variance.

  • Flag transition zones between forested and non-forested areas as high-uncertainty regions, enabling users to avoid over-trusting predictions in these challenging areas.

  • Handle extreme label sparsity (only a small fraction of pixels have ground-truth annotations) without degrading uncertainty calibration.

  • Support carbon accounting and biodiversity monitoring by providing not just point estimates but also confidence bounds, enabling risk-aware ecological decision-making.

Sources

Related papers