Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA

arXiv:2608.11271 · eess.IV, cs.LG · Submitted 2026-08-11 · Read on arXiv

Cédric Léonard, Francescopaolo Sica, Martin Schulz

Technical University of Munich · German Aerospace Center · University of the Bundeswehr Munich

eess.IV, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: Submitted to IEEE Transactions on Geoscience and Remote Sensing (TGRS). 11 pages, 8 figures, 4 tables

Code: https://github.com/CedricLeon/SAR_DDC_FPGA

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper presents the first deployment of a joint SAR Despeckling and Data Compression (DDC) framework on an embedded FPGA-based platform, using a ZCU102 board with AMD's Vitis AI overlay

Terminology

Summary

This paper presents the first deployment of a joint SAR Despeckling and Data Compression (DDC) framework on an embedded FPGA-based platform, using a ZCU102 board with AMD's Vitis AI overlay accelerator. The work bridges the gap between learned SAR compression methods and the strict power, compute, and operational constraints of spaceborne systems.

The authors introduce hardware-aware model adaptations to comply with the accelerator's fixed-point arithmetic (int8) and restricted operator set. Specifically, they replace GDN activation functions with ReLU, reshape transposed convolution kernels from (5×5, 2, 1) to (4×4, 1, 0), and partition the computational graph into DPU and CPU subgraphs—mapping neural transforms (ga, ha, hs, gs) onto the DPU while offloading entropy coding to the ARM CPU. This partition also enables parallelizing the real and imaginary passes through the main encoder and decoder across DPU cores.

The study evaluates four model topologies (FP, ResFP, SH, ResSH) across precision levels (float32 vs. int8) and across CPU, GPU, and FPGA platforms. Key findings include:

  1. Replacing GDN with ReLU improves rate-distortion performance: The ablation shows replacing GDN by ReLU improves the RD tradeoff by 1.67 ± 0.10 dB and 0.26 ± 0.07 bpp at the highest evaluated rate (λ = 1,000). This is unexpected since GDN was designed for natural-image compression, suggesting design principles established for compression of natural images do not necessarily transfer to SAR imagery.

  2. Residual blocks offer little benefit for ten times the compute: ResSH requires 156.68 GOPs versus 15.70 GOPs for SH, yet the improvement of ResSH over SH (0.75 ± 0.31 dB at λ = 1,000) is consistent but modest. The authors recommend the lightweight FP architecture for onboard deployment.

  3. Quantization effects: For all architectures, the FPGA int8 quantized model scores higher than the full-precision GPU baseline, by +0.50 (FP) to +1.55 dB (ResSH) at the highest rate in PSNR, though SSIM drops by ≈0.02 points and Edge Preservation Degree drops by 0.25–0.33. The authors attribute this to int8 reconstructions being systematically less bright than their float32 counterparts, matching more closely the despeckled reference across the dark majority of the scenes.

  4. Energy efficiency: The FPGA is significantly more energy-efficient, consuming 11–17× less energy per patch than the CPU and 2–6× less than the GPU. The ZCU102 draws 9.8–12.7 W of active power versus 108–126 W for the RTX A4000 GPU, comfortably within the 20–95 W budget typical of SmallSat payloads.

  5. Latency breakdown: FP has the lowest latency (22.0 ms per patch), with the log-normalization overhead on the un-vectorized ARM CPU reaching up to 41% of the compute time for FP. Parallelizing ga(Re2) and ga(Im2) saves 4.4 ms (42.0%) for ga, or 35.6 ms (48.7%) for ga with residual blocks.

  6. End-to-end projection: For a full TerraSAR-X SLC tile (21,036×29,828 resolution cells), compressing would take 11.3 minutes and 8.6 kJ with ResSH, or 4.0 minutes and 2.36 kJ with FP. The 2.51 GB tile shrinks to 123–132 MB (CR 19.0–20.5) at the highest rate, or 12.7–14.2 MB (CR 177–198) near the RD knee.

The paper identifies several limits: the method requires fully focused SLC images and produces non-complex despeckled images, necessitating onboard focusing; the flexible Vitis AI DPU is likely suboptimal for LIC; and future work could explore Quantization-Aware Training (QAT), a specific FPGA design, streaming pipelines, and task-specific decoders attached to the learned latent.

Improvements for AI systems

Improvements to AI Systems:

  1. Domain-Adaptive Activation Function Selection: Replace GDN with ReLU in compression autoencoders when processing SAR imagery, yielding +1.67 dB PSNR and −0.26 bpp at high rates. The improved system can achieve better rate-distortion tradeoffs for radar data without architectural redesign.

  2. Hardware-Aware Quantization Without Retraining: Use int8 fixed-point quantization directly on trained float32 models, which paradoxically improves PSNR by +0.50 to +1.55 dB for SAR despeckling (due to systematic brightness reduction matching dark reference scenes). The improved system can deploy on edge devices with lower power while gaining fidelity, though it must monitor SSIM (drop 0.02) and edge preservation (drop 0.25–0.33) for quality assurance.

  3. Operator-Set Constrained Kernel Reshaping: Convert transposed convolution kernels from (5×5, stride=2, pad=1) to (4×4, stride=1, pad=0) to fit DPU accelerators, enabling full neural transform execution on FPGA without precision loss. The improved system can run end-to-end learned compression on space-grade hardware with fixed operator libraries.

  4. Heterogeneous Graph Partitioning for Parallelism: Split computational graphs into DPU (neural transforms) and CPU (entropy coding) subgraphs, then parallelize real and imaginary passes across DPU cores—reducing encoder latency by 42–48.7%. The improved system can achieve 22.0 ms/patch latency on embedded platforms, enabling real-time onboard SAR compression.

  5. Energy-Aware Model Selection: Prefer lightweight architectures (FP, 15.7 GOPs) over residual-heavy ones (ResSH, 156.7 GOPs) when the RD gain is <1 dB, given 10× compute cost. The improved system can autonomously choose between fidelity and power budgets, operating at 9.8–12.7 W (vs. 108–126 W GPU), fitting SmallSat power envelopes.

  6. Quantization-Aware Brightness Calibration: Exploit the systematic int8 darkening as a feature—not a bug—by intentionally quantizing to match despeckled reference statistics. The improved system can produce visually cleaner outputs for dark-majority SAR scenes without extra post-processing.

  7. Streaming Pipeline Design: Offload log-normalization (up to 41% of compute time) to dedicated vectorized hardware or overlap with DPU execution, reducing end-to-end latency further. The improved system can process full TerraSAR-X tiles (21k×30k) in 4 minutes with FP, enabling rapid revisit rates for disaster monitoring.

  8. Task-Specific Latent Decoders: Attach lightweight task decoders (e.g., for target detection or classification) directly to the learned latent space, bypassing full image reconstruction. The improved system can perform onboard SAR analysis (despeckling + object recognition) in a single forward pass, saving bandwidth and compute.

What the Improved AI System Can Do:

  • Deploy learned SAR compression on a 9.8–12.7 W FPGA for SmallSats, compressing 2.51 GB tiles to 123–132 MB (CR 20) in 4 minutes, or to 12.7–14.2 MB (CR 190) near the RD knee.

  • Achieve higher PSNR than full-precision GPU baselines while using 2–6× less energy than GPU and 11–17× less than CPU.

  • Parallelize complex-valued SAR processing across dual DPU cores, enabling real-time despeckling and compression for streaming downlink.

  • Adapt activation functions and kernel shapes automatically based on sensor modality (SAR vs. optical) and hardware constraints, without retraining.

  • Provide a blueprint for other non-natural image domains (e.g., medical, hyperspectral) where GDN assumptions fail, by using ReLU + int8 quantization as a free fidelity boost.

Abstract

Next-generation Synthetic Aperture Radar (SAR) missions will generate data far faster than they can downlink, making onboard data reduction essential for near-real-time Earth observation. Learned Image Compression (LIC) offers better rate-distortion performance than handcrafted codecs used operationally today, and recent work shows that simultaneously despeckling and compressing SAR imagery enables better representation capacity while unlocking higher compression rates. These methods, however, have yet to be confronted with the strict power, compute, and operational constraints of spaceborne systems. In this work, we bridge this gap by deploying a joint SAR Despeckling and Data Compression (DDC) framework on an embedded ZCU102 FPGA-based platform, introducing model adaptations that respect the accelerator's fixed-point arithmetic and limited set of supported operations. We evaluate four model topologies across precision levels and across CPU, GPU, and FPGA platforms, revealing several findings with direct design implications. We find that replacing conventional GDN activation functions with plain ReLU improves quality on SAR, suggesting that design principles established for compression of natural images do not necessarily transfer to SAR imagery. In addition, we demonstrate that residual blocks offer little representational benefit for ten times the compute, and show that the FPGA is the most energy-efficient of the platforms tested. Together, these results set a functioning edge deployment workflow and an evidence-based starting point for onboard SAR compression. The code is available at https://github.com/CedricLeon/SAR DDC FPGA.

Sources

Related papers