Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

arXiv:2608.11770 · cs.CV, cs.DC, cs.LG · Submitted 2026-08-12 · Read on arXiv

Vaishnav Raju

Newspace Research and Technologies

cs.CV, cs.DC, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 14 pages, 17 figures, 5 tables. Submitted to Journal of Real-Time Image Processing

Code: https://github.com/NVIDIA-AI-IOT/jetson

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper presents a systematic five-step methodology for deploying custom classification models on NVIDIA Jetson Deep Learning Accelerator (DLA) cores with zero GPU fallback, enabling parallel

Terminology

Summary

This paper presents a systematic five-step methodology for deploying custom classification models on NVIDIA Jetson Deep Learning Accelerator (DLA) cores with zero GPU fallback, enabling parallel multi-model inference on a single edge SoC. The work targets the problem that "Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. The authors note that Modern edge SoCs pair GPUs with dedicated neural accelerators (NPUs, DLAs) capable of concurrent execution, yet deploying custom models on these accelerators remains impractical due to strict operator constraints, quantization incompatibilities, and an undocumented end-to-end pipeline."

The five-step methodology comprises: "(1) architecture adaptation, (2) manual dynamic range workaround to rescue TensorRT's implicit quantization (recovering 94.0% accuracy from implicit quantization's 75%) for rapid pipeline validation before explicit quantization, (3) quantization-aware training, (4) ONNX graph surgery for DLA compilation, and (5) a concurrent GPU-detection/DLA-classification inference pipeline. The paper documents nine engineering constraints with root-cause analysis and generalizable solutions."

Key technical findings include:

  • Operator replacements: Standard classification layers must be replaced with DLA-native equivalents: nn.Linear(C, num classes) → Conv2d(C, num classes, kernel size=1), AdaptiveAvgPool2d(1) → AvgPool2d(kernel size=k), and multi-head output concatenation must be replaced with fused single-head convolution where K heads of Conv2d(C→1) are fused into a single Conv2d(C→K) at export time.

  • Quantization failure: "Standard Post Training Quantization (PTQ) entropy calibration (ENTROPY CALIBRATION 2) produces significantly degraded accuracy for DLA INT8 deployment: 75% with ReLU6 and as low as 66% per-head with standard ReLU, both well below the FP32 baseline. The root cause is that entropy minimization selects a small clipping threshold T to minimize KL divergence between the FP32 and quantized distributions. Activations exceeding T saturate at INT8 value 127, and this saturation propagates through 34 layers of residual blocks."

  • Manual dynamic ranges: Structurally derived ranges recover accuracy to 94.0%: Input tensors: [−4, 4]... Intermediate tensors: [−8, 8]... Output tensors: [−1, 1]. ReLU6 is recommended because its known output bounds [0, 6] provide a structural anchor for INT8 ranges without calibration data.

  • QAT with percentile calibration: Percentile calibration at the 99.99th percentile produces the best results for this model class as the 0.01% of clipped values are transient residual-add spikes that carry no classification signal.

  • Graph surgery: "A custom graph surgery script traverses the ONNX graph, extracts per-tensor INT8 scales from Q/DQ node pairs, propagates through pass-through ops and residual connections, strips Q/DQ nodes, and produces a TRT calibration cache. The script requires onnxsim.simplify first to remove no-op Cast/Identity nodes [that] break graph traversal for scale propagation."

  • Pipeline architecture: The frame N−1 async dispatch design makes classification near-zero-overhead where the GPU processes detections on the current frame while the DLA classifies crops from the previous frame.

Validation results on a Jetson Orin NX with a dual-head person attribute classifier alongside a YOLOv11m detector show:

  • QAT DLA INT8 achieves 95.0% accuracy, matching or marginally exceeding the FP32 baseline of 94.5%, with zero GPU fallback.

  • PTQ manual ranges achieve 94.0% accuracy, within 1 pp of QAT while requiring no training time.

  • DLA compute latency is approximately 19.7 ms at batch size 16, with throughput around 50 QPS.

  • Pipeline throughput at 1080p: YOLO only (GPU) 13.3 FPS, YOLO + 1 clf sequential (GPU) 10.5 FPS, YOLO + 2 clf sequential (GPU) 9.6 FPS, YOLO + 1 clf parallel (GPU+DLA0) 12.5 FPS, YOLO + 2 clf parallel (GPU+DLA0+DLA1) 12.5 FPS.

  • At 720p: YOLO only (GPU) 27.5 FPS, YOLO + 1 clf sequential (GPU) 19.0 FPS, YOLO + 2 clf sequential (GPU) 15.3 FPS, YOLO + 1 clf parallel (GPU+DLA0) 25.0 FPS, YOLO + 2 clf parallel (GPU+DLA0+DLA1) 25.0 FPS.

The paper explains why DLA offloading achieves near-zero overhead while GPU multi-stream cannot: "CUDA multi-stream allows two kernels to be enqueued concurrently, but both compete for the same physical resources, namely streaming multiprocessors (SMs), L2 cache, and DRAM bandwidth... Any GPU cycles consumed by the classifier are cycles unavailable to the detector. In contrast, DLA is architecturally separate hardware with its own compute engines, local SRAM, and power domain... When DLA executes the classifier, the GPU's SMs, caches, and memory bandwidth are completely unaffected."

The ablation studies confirm: removing skip connection quantizer causes either FP16 on DLA: 16 Add nodes in FP16, 126.98 ms (6.4× vs production) or GPU fallback: 36 layers cascade to GPU (Add + conv2 + downsample), GPU occupied, pipeline parallelism lost; ReLU instead of ReLU6 causes QAT accuracy drops 95%→94% (−1pp), AUC −1.47pp/−1.37pp (Hd-A/Hd-B); removing explicit post backbone and post head quantizers causes All layers downstream of layer4 assigned FP16.

The methodology is claimed to be backbone-agnostic and detector-agnostic, generalizing to any detection-classification edge pipeline. The authors note that the benefit of DLA offloading scales with GPU model complexity, as heavier detectors leave even less GPU capacity for concurrent classifiers, making DLA offloading increasingly valuable. The paper concludes that this work fills a critical gap between NVIDIA's hardware capabilities and the practical engineering knowledge required to exploit them, enabling full utilization of the Jetson SoC's heterogeneous compute where Deploying classifiers on both DLA cores alongside GPU detection achieves identical throughput to a single-DLA configuration.

Improvements for AI systems

Improvements to AI Systems:

  1. Hierarchical Edge Inference with Zero-Contention Parallelism
  • Implement a dual-accelerator pipeline where a GPU runs detection (e.g., YOLO) on the current frame while DLA cores classify crops from the previous frame. This eliminates serial GPU bottlenecks, achieving 12.5 FPS at 1080p (vs. 9.6 FPS sequential) and 25.0 FPS at 720p (vs. 15.3 FPS sequential) for two classifiers, with no added latency.
  1. DLA-Native Model Architecture Adaptation
  • Automatically replace standard classification layers with DLA-compatible equivalents: nn.Linear → 1x1 Conv2d, AdaptiveAvgPool2d(1) → fixed-kernel AvgPool2d, and multi-head outputs → fused single-head convolution. This enables direct DLA compilation without GPU fallback, preserving full model accuracy.
  1. Rescue of Implicit Quantization via Manual Dynamic Ranges
  • Bypass TensorRT’s entropy calibration (which degrades accuracy to 75%) by setting structurally derived INT8 ranges: input [-4,4], intermediate [-8,8], output [-1,1], using ReLU6 activations as anchors. This recovers 94.0% accuracy (vs. FP32 94.5%) with zero training time, enabling rapid pipeline validation.
  1. Quantization-Aware Training with Percentile Calibration
  • Train with QAT using 99.99th percentile calibration to clip only transient residual-add spikes (which carry no classification signal). This achieves 95.0% INT8 accuracy, matching FP32, while avoiding accuracy loss from aggressive clipping.
  1. Automated ONNX Graph Surgery for DLA Deployment
  • Run onnxsim.simplify to remove no-op Cast/Identity nodes, then traverse the graph to extract per-tensor scales from Q/DQ pairs, propagate through pass-through ops and residual connections, strip Q/DQ nodes, and generate a TRT calibration cache. This makes DLA compilation reproducible and eliminates manual scale tuning.
  1. Skip-Connection Quantizer Placement
  • Explicitly insert quantizers on skip connections and post-backbone/head layers to prevent FP16 fallback on DLA. This avoids 6.4× latency increase (126.98 ms vs. 19.7 ms) and keeps all 34 residual layers in INT8, preserving pipeline parallelism.
  1. Heterogeneous Compute Utilization for Scaling
  • Deploy classifiers on both DLA0 and DLA1 alongside GPU detection to achieve identical throughput to single-DLA (12.5 FPS at 1080p), effectively doubling classifier capacity without sacrificing detection speed—critical for heavier detectors or multi-attribute analysis.
  1. Backbone-Agnostic Deployment Framework
  • Generalize the five-step methodology to any detection-classification pipeline (e.g., person attribute, vehicle color, animal species) on Jetson devices, enabling rapid porting of new models without re-engineering the quantization or compilation workflow.

What the Improved AI System Can Do:

  • Run real-time hierarchical vision (detection + multiple fine-grained classifiers) on a single edge SoC at full accuracy (≥95% INT8) with zero GPU fallback.

  • Achieve near-zero-overhead classification (19.7 ms per batch of 16) by offloading to DLA, freeing GPU for heavier detectors.

  • Automatically adapt, quantize, and compile custom models for DLA in minutes, reducing deployment time from weeks to hours.

  • Scale to multi-classifier pipelines (e.g., person + clothing + action) without throughput loss, enabling richer edge analytics for surveillance, drones, and autonomous vehicles.

Sources

Related papers