Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines
Vaishnav Raju
Newspace Research and Technologies
cs.CV, cs.DC, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 14 pages, 17 figures, 5 tables. Submitted to Journal of Real-Time Image Processing
Code: https://github.com/NVIDIA-AI-IOT/jetson
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This paper presents a systematic five-step methodology for deploying custom classification models on NVIDIA Jetson Deep Learning Accelerator (DLA) cores with zero GPU fallback, enabling parallel
Terminology
Summary
This paper presents a systematic five-step methodology for deploying custom classification models on NVIDIA Jetson Deep Learning Accelerator (DLA) cores with zero GPU fallback, enabling parallel multi-model inference on a single edge SoC. The work targets the problem that "Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. The authors note that
Modern edge SoCs pair GPUs with dedicated neural accelerators (NPUs, DLAs) capable of concurrent execution, yet deploying custom models on these accelerators remains impractical due to strict operator constraints, quantization incompatibilities, and an undocumented end-to-end pipeline."
The five-step methodology comprises: "(1) architecture adaptation, (2) manual dynamic range workaround to rescue TensorRT's implicit quantization (recovering 94.0% accuracy from implicit quantization's 75%) for rapid pipeline validation before explicit quantization, (3) quantization-aware training, (4) ONNX graph surgery for DLA compilation, and (5) a concurrent GPU-detection/DLA-classification inference pipeline. The paper documents
nine engineering constraints with root-cause analysis and generalizable solutions."
Key technical findings include:
-
Operator replacements: Standard classification layers must be replaced with DLA-native equivalents:
nn.Linear(C, num classes) → Conv2d(C, num classes, kernel size=1)
,AdaptiveAvgPool2d(1) → AvgPool2d(kernel size=k)
, and multi-head output concatenation must be replaced withfused single-head convolution
where K heads of Conv2d(C→1) are fused into a single Conv2d(C→K) at export time. -
Quantization failure: "Standard Post Training Quantization (PTQ) entropy calibration (ENTROPY CALIBRATION 2) produces significantly degraded accuracy for DLA INT8 deployment: 75% with ReLU6 and as low as 66% per-head with standard ReLU, both well below the FP32 baseline.
The root cause is that
entropy minimization selects a small clipping threshold T to minimize KL divergence between the FP32 and quantized distributions. Activations exceeding T saturate at INT8 value 127, and this saturation propagates through 34 layers of residual blocks." -
Manual dynamic ranges: Structurally derived ranges recover accuracy to 94.0%:
Input tensors: [−4, 4]... Intermediate tensors: [−8, 8]... Output tensors: [−1, 1].
ReLU6 is recommended becauseits known output bounds [0, 6] provide a structural anchor for INT8 ranges without calibration data.
-
QAT with percentile calibration:
Percentile calibration at the 99.99th percentile produces the best results for this model class as the 0.01% of clipped values are transient residual-add spikes that carry no classification signal.
-
Graph surgery: "A custom graph surgery script traverses the ONNX graph, extracts per-tensor INT8 scales from Q/DQ node pairs, propagates through pass-through ops and residual connections, strips Q/DQ nodes, and produces a TRT calibration cache.
The script requires onnxsim.simplify first to remove
no-op Cast/Identity nodes [that] break graph traversal for scale propagation." -
Pipeline architecture: The frame N−1 async dispatch design
makes classification near-zero-overhead
wherethe GPU processes detections on the current frame while the DLA classifies crops from the previous frame.
Validation results on a Jetson Orin NX with a dual-head person attribute classifier alongside a YOLOv11m detector show:
-
QAT DLA INT8 achieves 95.0% accuracy,
matching or marginally exceeding the FP32 baseline
of 94.5%, with zero GPU fallback. -
PTQ manual ranges achieve 94.0% accuracy,
within 1 pp of QAT while requiring no training time.
-
DLA compute latency is approximately 19.7 ms at batch size 16, with throughput around 50 QPS.
-
Pipeline throughput at 1080p:
YOLO only (GPU) 13.3 FPS, YOLO + 1 clf sequential (GPU) 10.5 FPS, YOLO + 2 clf sequential (GPU) 9.6 FPS, YOLO + 1 clf parallel (GPU+DLA0) 12.5 FPS, YOLO + 2 clf parallel (GPU+DLA0+DLA1) 12.5 FPS.
-
At 720p:
YOLO only (GPU) 27.5 FPS, YOLO + 1 clf sequential (GPU) 19.0 FPS, YOLO + 2 clf sequential (GPU) 15.3 FPS, YOLO + 1 clf parallel (GPU+DLA0) 25.0 FPS, YOLO + 2 clf parallel (GPU+DLA0+DLA1) 25.0 FPS.
The paper explains why DLA offloading achieves near-zero overhead while GPU multi-stream cannot: "CUDA multi-stream allows two kernels to be enqueued concurrently, but both compete for the same physical resources, namely streaming multiprocessors (SMs), L2 cache, and DRAM bandwidth... Any GPU cycles consumed by the classifier are cycles unavailable to the detector. In contrast,
DLA is architecturally separate hardware with its own compute engines, local SRAM, and power domain... When DLA executes the classifier, the GPU's SMs, caches, and memory bandwidth are completely unaffected."
The ablation studies confirm: removing skip connection quantizer causes either FP16 on DLA: 16 Add nodes in FP16, 126.98 ms (6.4× vs production)
or GPU fallback: 36 layers cascade to GPU (Add + conv2 + downsample), GPU occupied, pipeline parallelism lost
; ReLU instead of ReLU6 causes QAT accuracy drops 95%→94% (−1pp), AUC −1.47pp/−1.37pp (Hd-A/Hd-B)
; removing explicit post backbone and post head quantizers causes All layers downstream of layer4 assigned FP16.
The methodology is claimed to be backbone-agnostic and detector-agnostic,
generalizing to any detection-classification edge pipeline.
The authors note that the benefit of DLA offloading scales with GPU model complexity, as heavier detectors leave even less GPU capacity for concurrent classifiers, making DLA offloading increasingly valuable.
The paper concludes that this work fills a critical gap between NVIDIA's hardware capabilities and the practical engineering knowledge required to exploit them,
enabling full utilization of the Jetson SoC's heterogeneous compute
where Deploying classifiers on both DLA cores alongside GPU detection achieves identical throughput to a single-DLA configuration.
Improvements for AI systems
Improvements to AI Systems:
- Hierarchical Edge Inference with Zero-Contention Parallelism
- Implement a dual-accelerator pipeline where a GPU runs detection (e.g., YOLO) on the current frame while DLA cores classify crops from the previous frame. This eliminates serial GPU bottlenecks, achieving 12.5 FPS at 1080p (vs. 9.6 FPS sequential) and 25.0 FPS at 720p (vs. 15.3 FPS sequential) for two classifiers, with no added latency.
- DLA-Native Model Architecture Adaptation
- Automatically replace standard classification layers with DLA-compatible equivalents:
nn.Linear→1x1 Conv2d,AdaptiveAvgPool2d(1)→ fixed-kernelAvgPool2d, and multi-head outputs → fused single-head convolution. This enables direct DLA compilation without GPU fallback, preserving full model accuracy.
- Rescue of Implicit Quantization via Manual Dynamic Ranges
- Bypass TensorRT’s entropy calibration (which degrades accuracy to 75%) by setting structurally derived INT8 ranges: input
[-4,4], intermediate[-8,8], output[-1,1], using ReLU6 activations as anchors. This recovers 94.0% accuracy (vs. FP32 94.5%) with zero training time, enabling rapid pipeline validation.
- Quantization-Aware Training with Percentile Calibration
- Train with QAT using 99.99th percentile calibration to clip only transient residual-add spikes (which carry no classification signal). This achieves 95.0% INT8 accuracy, matching FP32, while avoiding accuracy loss from aggressive clipping.
- Automated ONNX Graph Surgery for DLA Deployment
- Run
onnxsim.simplifyto remove no-op Cast/Identity nodes, then traverse the graph to extract per-tensor scales from Q/DQ pairs, propagate through pass-through ops and residual connections, strip Q/DQ nodes, and generate a TRT calibration cache. This makes DLA compilation reproducible and eliminates manual scale tuning.
- Skip-Connection Quantizer Placement
- Explicitly insert quantizers on skip connections and post-backbone/head layers to prevent FP16 fallback on DLA. This avoids 6.4× latency increase (126.98 ms vs. 19.7 ms) and keeps all 34 residual layers in INT8, preserving pipeline parallelism.
- Heterogeneous Compute Utilization for Scaling
- Deploy classifiers on both DLA0 and DLA1 alongside GPU detection to achieve identical throughput to single-DLA (12.5 FPS at 1080p), effectively doubling classifier capacity without sacrificing detection speed—critical for heavier detectors or multi-attribute analysis.
- Backbone-Agnostic Deployment Framework
- Generalize the five-step methodology to any detection-classification pipeline (e.g., person attribute, vehicle color, animal species) on Jetson devices, enabling rapid porting of new models without re-engineering the quantization or compilation workflow.
What the Improved AI System Can Do:
-
Run real-time hierarchical vision (detection + multiple fine-grained classifiers) on a single edge SoC at full accuracy (≥95% INT8) with zero GPU fallback.
-
Achieve near-zero-overhead classification (19.7 ms per batch of 16) by offloading to DLA, freeing GPU for heavier detectors.
-
Automatically adapt, quantize, and compile custom models for DLA in minutes, reducing deployment time from weeks to hours.
-
Scale to multi-classifier pipelines (e.g., person + clothing + action) without throughput loss, enabling richer edge analytics for surveillance, drones, and autonomous vehicles.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models