Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study
summary
The gist
Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested
In short
This study measured how long it takes for a sensor interrupt to trigger a decision on an NVIDIA Jetson Orin Nano using either a host computer or an onboard Machine Learning Core (MLC). The results show that the host inference method has lower latency than the MLC pipeline under all tested conditions. The primary bottleneck is not the ML core's calculation, but rather the overhead of reading data via I2C protocols.
Key concepts
- Wire-Level Latency
- This measures the actual time delay from a physical signal event (like an interrupt edge) to a corresponding electrical change on a wire (like a GPIO pin). It is measured using high-resolution logic analyzers to bypass software timing uncertainties, providing an accurate picture of hardware performance.
- MLC Pipeline
- This refers to the onboard Machine Learning Core running on the sensor itself. The study tested different ways this core operates, specifically comparing a standard version that performs three I2C reads per interrupt against a binary version that skips the read entirely to isolate system overhead.
- I2C Read Protocol Overhead
- The time spent communicating with the sensor via the I2C bus is identified as the main source of delay. The study found that even when comparing different ML processing methods, the extra time required for these three-transaction I2C reads significantly slows down the overall decision process.
- Interrupt-to-Decision Latency
- This is the total time elapsed from when a sensor generates an interrupt signal to when the host system's output GPIO line changes its state based on that interrupt. This metric was used to compare how quickly different processing methods (host versus MLC) can react to incoming sensor data.
Terminology used across episodes
This episode discusses
- Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study · Paper Radio
- Event-Driven On-Sensor Locomotion Mode Recognition Using a Shank-Mounted IMU with Embedded Machine Learning for Exoskeleton Control
- Per-Platform GPIO Overhead in Hardware-Validated Edge ML Inference Timing
The paper
Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study · Read on arXiv
Akul Swami, Dnyaneshwar Sonawane
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano".
Rosa: Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested conditions.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: To summarize, the core thesis of this paper is that while embedded Machine Learning Cores are touted for low latency, they haven't really been measured at the wire level concerning the total interrupt-to-decision time. Akul Swami and Dnyaneshwar Sonawane used a Saleae Logic Pro eight logic analyzer on an NVIDIA Jetson Orin Nano to measure this latency across three distinct pipelines under three different stress conditions: idle, I2C bus contention, and CPU saturation.
Dev: They are testing the interrupt-to-decision latency specifically, meaning they are tracking the time from when a sensor generates an INT1 edge all the way to when the host side finally executes a decision GPIO.
Rosa: The results show that across all tested conditions, this study claims that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core pipelines.
Dev: This finding is significant because it suggests that the dominant source of delay isn't necessarily the silicon performing the classification itself, but rather the overhead introduced by the I2C read protocol in some of those MLC scenarios.
Rosa: They set up three specific comparison pipelines: a host-side decision-tree classifier, a standard MLC bank-switch read protocol involving three I2C transactions, and an MLC binary variant that omits the I2C read entirely.
Dev: The comparison wasn't just about raw speed; they also looked at how those latencies shifted when the system was under stress, specifically showing that the host pipeline maintained a significant advantage even when there was heavy I2C bus contention.
Rosa: They also found some structural observations regarding the MLC operation, documenting a reproducible seven hundred six point five ms decision cadence and noting that for the MLC pipelines at idle, they observed multimodal distributions where the upper mode was actually more representative of worst-case timing specifications than just the median.
Dev: This paper is important because it addresses that gap between functional correctness—the AI is right—and timing correctness—the AI gives us an answer fast enough to act on it in real-time systems, which is a major concern for autonomous applications.
Taro: From my perspective, this study really gets to the heart of the issue we face in building reliable autonomous systems; it’s not just about having a smart sensor but ensuring that intelligence is delivered with the required timing guarantees when the environment starts throwing noise at you.
Rosa: It sets a clear benchmark by quantifying this wire-level delay for different operational modes, which helps developers understand exactly where their latency budget needs to be focused.
Conclusion: Rosa: Looking at the full context of this paper, Akul Swami and Dnyaneshwar Sonawane’s work on "Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study" really zeroes in on that specific timing gap we talked about earlier.
Dev: They are essentially confirming that the standard way some MLCs handle data delivery through those I2C transactions introduces measurable delay compared to a direct host-side approach, especially when the system is running under load or contention.
Rosa: The implication here is practical: if you're designing a wearable exoskeleton control system or something safety-critical, you can't just trust the functional result of an embedded AI without rigorously measuring this wire-level latency because it dictates whether the system can physically react in time.
Dev: Exactly, and their findings about how the MLC binary variant isolates that kernel/gpiod latency floor really shows us that when we strip away the ML computation itself, we see exactly what’s slowing down our decision path on this hardware.
Taro: So, for autonomy research, this means if you're relying on sensor fusion where decisions need to be made in milliseconds, you have to account for these I2C protocol overheads because they can unexpectedly inflate the total time it takes for a reaction.
Rosa: And that’s why their finding that the three-transaction I2C read protocol amplifies contention penalties is so important—it tells us exactly what kind of operational stress could push a system over its timing budget.
Dev: It's a very concrete piece of data because they didn't just talk about theory; they used external timestamped measurements to prove that the host pipeline is faster in many situations, which helps us build more realistic performance models for deployment.
Taro: I think the broader impact is forcing the industry to move beyond just reporting ML accuracy and start demanding latency specifications that account for these underlying hardware communication layers when building autonomous agents.
Rosa: It moves the focus from "how good is your classifier?" to "can your entire sensing-to-action loop meet its time constraints?"
Dev: So, in simple terms, the study on this paper confirms that for many scenarios on the Jetson Orin Nano, running inference directly on the host side beats the standard embedded protocol because of communication overhead.
Taro: That really hammers home that when you're building things that interact with a physical world, you need to be meticulous about those communication details, not just the high-level algorithms.
Rosa: It’s about understanding how much time is spent waiting for a signal to travel and be processed by the I2C stack versus the actual intelligence calculation itself.
Dev: Precisely. The paper gives us a concrete measure of that waiting time under stress, which means we can finally start designing systems that are timing-aware from the beginning rather than trying to patch performance issues later on when things go wrong in deployment.
Taro: That seems like a very useful contribution for anyone working on real-time AI where failure modes involve missed deadlines, and it helps us predict those failure modes more accurately.
Rosa: It gives us a solid foundation for performance requirements that aren't just theoretical ideals but are backed by measured wire-level data showing the actual behavior under various operating conditions on this specific hardware setup.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications