Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation

arXiv:2608.11810 · cs.CV, cs.LG · Submitted 2026-08-12 · Read on arXiv

Hyewook Kim, Byul Kang, Seokbin Yoon, Keumjin Lee

Korea Aerospace Research Institute · Korea Aerospace University

cs.CV, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 12 pages, 6 figures. Submitted to Elsevier

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper investigates whether radar display imagery, despite its atypical visual characteristics, is a viable input format for deep learning vision models in the context of air traffic complexity

Terminology

Summary

This paper investigates whether radar display imagery, despite its atypical visual characteristics, is a viable input format for deep learning vision models in the context of air traffic complexity estimation. The authors identify two fundamental characteristics that distinguish radar imagery from natural images: (C1) extreme sparsity and self-similarity, where radar images consist of a nearly uniform black background with a few visually identical aircraft blobs, making different images highly similar in appearance (e.g., two different traffic situations may differ in only 1.05% of pixels, whereas cat and dog images differ in 99.66% of pixels); and (C2) high sensitivity of complexity to few-pixel changes, where displacing, adding, or removing a single aircraft can substantially change sector-level complexity, which is the opposite of what modern vision models are designed for (they are typically invariant to small pixel perturbations).

To test feasibility, the authors generate 100,000 synthetic radar images using the BlueSky ATM simulator across three route configurations (single, crossing, and merging), with aircraft counts ranging from 2 to 10. The input representation consists of the original radar image as a position channel, supplemented by five additional channels encoding heading, speed, current altitude, requested altitude, and vertical speed at the pixel locations of each aircraft blob, yielding a six-channel input tensor. No handcrafted features such as pairwise distances or conflict indicators are computed. Labels are generated using the intrinsic complexity metric of Delahaye and Puechmorel (2000), extended to three dimensions, producing four sector-level components: density, convergence, divergence, and insensitivity to control actions.

The model architecture is a Vision Transformer (ViT-Base) with 12 encoder layers, embedding dimension 768, 12 attention heads, and patch size 24 (yielding 121 patch tokens per image). Four dedicated component query tokens (CQTs), one for each complexity component, are prepended to the patch sequence, each with its own MLP regression head. A key methodological contribution is a differentiated attention masking strategy: patch-to-patch attention is left unmasked so that empty background patches can carry spatial relations among aircraft, while each CQT attends only to aircraft-bearing patches to avoid diluting the few informative tokens among many empty ones.

Results show strong prediction performance across all four components, with overall R squared > 0.96 (density: 0.9936, convergence: 0.9846, divergence: 0.9677, insensitivity: 0.9906) and small absolute errors (MAE 0.0148–0.0271). Performance varies by route configuration: single-route scenarios yield lower R squared for convergence and divergence (0.782 and 0.774) because label values cluster near zero with little variance, though absolute errors remain smallest there (MAE 0.0126 and 0.0161). The authors interpret low R squared in these cases as a property of the evaluation metric rather than model limitation.

A one-aircraft-removal perturbation study evaluates whether the model responds proportionally to the removed aircraft's actual contribution to sector complexity. For each test image, one aircraft is randomly removed, and the change in ground-truth complexity C is compared to the change in model prediction. Results show strong proportionality for density, convergence, and divergence (R squared at least 0.92), with density most accurate (slope = 1.10, R squared = 0.94, RMSE = 0.02), while convergence and divergence slightly underestimate large changes (slopes 0.89 and 0.84). This demonstrates the model has learned relational contributions rather than merely counting aircraft, since every removal changes roughly the same number of pixels. The insensitivity component is excluded from this analysis because removing an aircraft changes the set of converging pairs, making before- and after-removal values incomparable.

The authors conclude that radar imagery is a viable data format for air traffic complexity modeling despite its atypical characteristics, clearing the ground for the hypothesis that vision models operating on the same visual information available to controllers can learn to model controller-perceived complexity. Limitations include the use of synthetic scenarios with fixed route configurations, one-directional traffic flows that leave divergence values clustered near zero, single-snapshot inputs without temporal context, and geometry-based rather than perception-based reference labels. Future work includes testing on real radar data, incorporating bi-directional flows, adding temporal context, and relating image-based estimates to controller-rated complexity.

Improvements for AI systems

Improvements to AI systems:

  1. Sparse-Input Vision Transformers with Differentiated Attention Masking
  • Modify ViT architectures to include query tokens that attend only to non-empty patches, while preserving unmasked patch-to-patch attention for spatial context. This prevents dilution of informative signals in extremely sparse images (e.g., <2% non-zero pixels).

  • The improved system can process any sparse visual input (e.g., satellite imagery, medical scans with few lesions, industrial defect maps) without performance collapse, achieving high regression accuracy even when pixel differences between classes are <2%.

  1. Perturbation-Proportional Prediction Heads
  • Add a training objective or post-hoc calibration that enforces linear proportionality between input changes (e.g., removing one object) and output changes, using slope and R2 metrics as regularizers.

  • The improved system can provide interpretable, causally consistent predictions: for any single-object removal, the predicted complexity shift matches the true shift within 10% error (slope 0.89–1.10), enabling reliable what-if analysis in safety-critical domains (e.g., air traffic control, power grid load, autonomous driving risk).

  1. Multi-Channel Positional Encoding for Non-Image Data
  • Extend the six-channel input (position + heading, speed, altitudes, vertical speed) to a general framework where any sparse spatial data is augmented with per-pixel attribute channels, without handcrafted features.

  • The improved system can directly ingest raw sensor feeds (radar, LiDAR, ultrasound) and learn relational dynamics (convergence, divergence) from pixel-level attributes alone, eliminating feature engineering pipelines and improving generalization to unseen configurations.

  1. Component-Specific Query Tokens for Multi-Output Regression
  • Use separate learnable query tokens per output component, each with its own attention mask and regression head, to decouple correlated targets.

  • The improved system can simultaneously predict multiple interdependent metrics (e.g., density, convergence, divergence, insensitivity) with R2 > 0.96, even when some targets have near-zero variance in certain regimes, and can flag when low R2 is due to label distribution rather than model failure.

  1. Robustness to Extreme Self-Similarity
  • Train on synthetic data with controlled route configurations and aircraft counts (2–10) to force the model to learn fine-grained differences (e.g., 1.05% pixel difference) rather than coarse visual features.

  • The improved system can distinguish between scenarios that differ by a single object’s position or presence, achieving 0.94 R2 on perturbation impact prediction, and can generalize to real-world sparse imagery with minimal fine-tuning.

  1. Sensitivity-Aware Loss Functions
  • Replace standard MSE with a loss that weights errors proportionally to the true complexity change (e.g., higher weight for high-impact perturbations), based on the finding that few-pixel changes can drastically alter output.

  • The improved system will prioritize accuracy on rare, high-stakes events (e.g., near-miss aircraft) rather than average cases, reducing catastrophic underprediction in safety applications.


What the improved AI system can do:

  • Real-time air traffic complexity estimation from raw radar screens, with per-component breakdowns (density, convergence, divergence, insensitivity) and causal explanations for each aircraft’s contribution, enabling controllers to test what if I remove/reroute this flight? with 94% accuracy.

  • Transfer to other sparse sensor domains (e.g., maritime radar, weather radar, autonomous vehicle occupancy grids) where inputs are mostly empty but single-object changes are critical, without retraining from scratch.

  • Provide uncertainty-aware predictions that distinguish low R2 due to label clustering (e.g., near-zero complexity) from genuine model error, allowing operators to trust outputs in both routine and extreme scenarios.

  • Scale to 100,000+ synthetic scenarios for stress-testing and adversarial robustness, while maintaining computational efficiency via sparse attention (only 121 patches, 12 layers).

Related papers