Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration

arXiv:2608.11638 · cs.LG, cs.CV · Submitted 2026-08-12 · Read on arXiv

Pann Thinzar Seint, Bryan Atwood, Subas Chhatkuli

DAI Labs

cs.LG, cs.CV

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper presents an operational framework for wall-to-wall above-ground biomass (AGB) estimation that combines a single globally trained convolutional neural network (CNN) with a lightweight

Terminology

Summary

The paper presents an operational framework for wall-to-wall above-ground biomass (AGB) estimation that combines a single globally trained convolutional neural network (CNN) with a lightweight field-calibration workflow. The global model integrates multi-sensor data—Sentinel-2 optical, Sentinel-1 C-band SAR, ALOS-2 PALSAR-2 L-band SAR, and Copernicus DEM—all harmonized onto a shared 10 m grid. It is trained once against GEDI Level-4A biomass reference data spanning multiple regions and both wet and dry seasons. The dual-season training design exposes the model to the full annual range of canopy greenness, leaf-on/leaf-off structure, and soil-moisture backscatter conditions, so it learns persistent woody-structure signals rather than single-date appearances. The CNN uses a two-stage convolutional process with adaptive pooling and a softplus activation, processing 5 × 5 × 24 patches (25-pixel neighborhoods across 24 bands). The model was trained on 281,872 patches and evaluated on 74,388 validation patches, using a hybrid loss combining log-domain SmoothL1 with RMSE to handle the skewed biomass distribution. On held-out validation, the global GEDI-based model attains R2 ≈ 0.78 and RMSE ≈ 22 Mg/ha (specifically R2 = 0.78, RMSE = 22.86 Mg/ha, bias = -1.09 Mg/ha, residual standard deviation = 22.83 Mg/ha).

The field-calibration workflow is applied per region using a minimum of 50 field plots per site, collected via nested circular plots in a "T" configuration (central plot with north, east, and south plots each 150 m from center), with trees recorded in nested concentric subplots by DBH thresholds. AGB is calculated using the Kachamba et al. (2016) allometric equation developed for miombo woodlands of Malawi. The calibration pipeline has two sequential phases: first, a Random Forest ensemble fine-tuning step (400 trees, squared-error criterion, max depth of five, square-root feature selection, 10-fold cross-validation with 70/30 train-validation splits per fold, using 7×7 moving window neighborhood features) that models non-linear regional variations; second, a third-order polynomial bias correction (AGBcal = a·AGB3pred + b·AGB2pred + c·AGBpred + d) fitted on 70% of field plots and validated on 30%, eliminating systematic scale and offset errors.

Results show that the uncalibrated global model performs poorly against local field plots—for Perekezi (2025), R2 = 0.11 and RMSE = 33.58 Mg/ha. After field calibration, performance improves dramatically: Perekezi (2025) reaches R2 = 0.82 and RMSE = 15.00 Mg/ha; Perekezi (2020) reaches R2 = 0.83 and RMSE = 15.51 Mg/ha; Ntchisi (2022) reaches R2 = 0.80 and RMSE = 12.82 Mg/ha; Dzalanyama (2020) reaches R2 = 0.82 and RMSE = 12.55 Mg/ha. The calibrated framework outperforms both the uncalibrated global model and the ESA CCI Biomass product when validated against field plots. Incorporating ALOS PALSAR-2 data in calibration improves accuracy—at 10 m resolution, R2 rises from 0.79 to 0.82 for Perekezi 2025, with a 6 m variant yielding further improvement.

The paper discusses applications in carbon accounting and climate policy (MRV protocols, REDD+, national GHG inventories), forest management and conservation, ecosystem and biodiversity assessments, wildfire modeling, and agricultural/infrastructure uses. The authors acknowledge limitations including GEDI footprint sampling biases across biomes and slopes, sensor saturation in high-biomass dense-canopy forests, and dependence of final accuracy on the spatial distribution of local field plots. The conclusion states that the localized calibration phase elevates validation performance from an uncalibrated R2 of 0.11 up to 0.82 while reducing error to just 15 Mg/ha, delivering an operational basis for regional biomass monitoring that is affordable in data requirements and directly serviceable to forest carbon accounting and climate mitigation.

Improvements for AI systems

Improvements to AI Systems:

  1. Dual-season training strategy – Train the model on data from both wet and dry seasons to learn invariant woody-structure signals, reducing sensitivity to phenological state and soil moisture. This improves generalization across temporal conditions.

  2. Hybrid loss function (log-domain SmoothL1 + RMSE) – Use a composite loss that balances absolute error in log space (for skewed distributions) with root-mean-square error (for large-magnitude accuracy). This yields better calibration for low- and high-biomass extremes.

  3. Two-stage convolutional architecture with adaptive pooling – Implement a CNN that processes 5×5×24 patches (neighborhood + multi-sensor bands) through two convolutional stages, then applies adaptive pooling to capture scale-invariant features. This reduces overfitting to local texture and improves transferability.

  4. Lightweight field-calibration pipeline – Integrate a two-phase post-processing module: (a) Random Forest fine-tuning on 7×7 neighborhood features with 10-fold cross-validation, then (b) a third-order polynomial bias correction fitted on 70% of field plots. This corrects systematic scale/offset errors without retraining the global model.

  5. Multi-sensor harmonization on a shared grid – Fuse Sentinel-2 optical, Sentinel-1 C-band SAR, ALOS-2 L-band SAR, and DEM into a single 24-band input. This improves robustness to cloud cover, canopy density, and terrain effects, enabling wall-to-wall mapping.

  6. Minimum-field-plot efficiency – Use only 50 field plots per site for calibration, with a nested "T" plot design and DBH-threshold subplots. This reduces data collection cost while achieving high accuracy, making the system operational in data-scarce regions.


What the Improved AI System Can Do:

  • Produce wall-to-wall above-ground biomass maps at 10 m resolution with R2 ≈ 0.80–0.83 and RMSE ≈ 12–15 Mg/ha after local calibration, even when the global model alone performs poorly (R2 = 0.11).

  • Operate across diverse ecosystems and seasons without retraining, by leveraging dual-season training and multi-sensor fusion to ignore leaf-on/off and soil-moisture artifacts.

  • Self-correct regional biases using a small set of field plots, automatically adjusting for allometric differences, sensor saturation, and local forest structure.

  • Generate carbon stock estimates for MRV, REDD+, and national GHG inventories with quantified uncertainty, directly supporting climate policy and carbon credit verification.

  • Detect biomass changes over time (e.g., 2020 vs. 2025) with consistent accuracy, enabling deforestation, degradation, and regrowth monitoring.

  • Provide high-resolution inputs for wildfire fuel load modeling, biodiversity habitat assessment, and sustainable forest management at scales previously infeasible with field surveys alone.

  • Extend to other regions with minimal effort – only 50 plots per new site are needed, making the system affordable for developing nations and remote areas.

Abstract

Spatially continuous quantification of forest above-ground biomass (AGB) is what makes carbon accounting credible and mitigation strategies actionable. While field inventories provide high localized accuracy, they are spatially sparse; conversely, spaceborne LiDAR from the Global Ecosystem Dynamics Investigation (GEDI) offers broad biomass samples but lacks spatial continuity and systematic underestimation of high-biomass forests. This paper presents an operational framework centered on a single globally trained convolutional neural network (CNN) that is seamlessly adapted to each new landscape through a lightweight empirical field-calibration workflow. The global model combines optical (Sentinel-2), C-band SAR (Sentinel-1), L-band SAR (ALOS-2 PALSAR-2), and terrain (DEM) data. It is trained once against GEDI Level-4A biomass reference data spanning multiple regions and both wet and dry seasons so that it learns the persistent woody-structure rather than a single-date appearance. To avoid retraining for every landscape, the framework applies a small number of local field plots to fit a scale-and-bias correction that aligns the global prediction with ground truth in each region. The pipeline harmonizes sensor data onto a shared 10 m grid, derives vegetation indices and polarimetric ratios, computes per-band normalization stats, and trains the CNN with a hybrid log-domain SmoothL1 with RMSE loss for skewed biomass distribution. On held-out validation the global GEDI-based model achieved R squared approximately 0.78 and RMSE approximately 22 Mg/ha. A subsequent field calibration combining Random Forest fine-tuning under a 10-fold cross-validation eliminates localized regional biases. This improves local validation performance to R squared approximately 0.82 and reduces RMSE to approximately 15 Mg/ha, outperforming both the uncalibrated global model and the ESA CCI Biomass product against field plots.

Related papers