Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano".
Rosa: Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested conditions.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: To summarize, the core thesis of this paper is that while embedded Machine Learning Cores are touted for low latency, they haven't really been measured at the wire level concerning the total interrupt-to-decision time. Akul Swami and Dnyaneshwar Sonawane used a Saleae Logic Pro eight logic analyzer on an NVIDIA Jetson Orin Nano to measure this latency across three distinct pipelines under three different stress conditions: idle, I2C bus contention, and CPU saturation.
Dev: They are testing the interrupt-to-decision latency specifically, meaning they are tracking the time from when a sensor generates an INT1 edge all the way to when the host side finally executes a decision GPIO.
Rosa: The results show that across all tested conditions, this study claims that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core pipelines.
Dev: This finding is significant because it suggests that the dominant source of delay isn't necessarily the silicon performing the classification itself, but rather the overhead introduced by the I2C read protocol in some of those MLC scenarios.
Rosa: They set up three specific comparison pipelines: a host-side decision-tree classifier, a standard MLC bank-switch read protocol involving three I2C transactions, and an MLC binary variant that omits the I2C read entirely.
Dev: The comparison wasn't just about raw speed; they also looked at how those latencies shifted when the system was under stress, specifically showing that the host pipeline maintained a significant advantage even when there was heavy I2C bus contention.
Rosa: They also found some structural observations regarding the MLC operation, documenting a reproducible seven hundred six point five ms decision cadence and noting that for the MLC pipelines at idle, they observed multimodal distributions where the upper mode was actually more representative of worst-case timing specifications than just the median.
Dev: This paper is important because it addresses that gap between functional correctness—the AI is right—and timing correctness—the AI gives us an answer fast enough to act on it in real-time systems, which is a major concern for autonomous applications.
Taro: From my perspective, this study really gets to the heart of the issue we face in building reliable autonomous systems; it’s not just about having a smart sensor but ensuring that intelligence is delivered with the required timing guarantees when the environment starts throwing noise at you.
Rosa: It sets a clear benchmark by quantifying this wire-level delay for different operational modes, which helps developers understand exactly where their latency budget needs to be focused.
Conclusion: Rosa: Looking at the full context of this paper, Akul Swami and Dnyaneshwar Sonawane’s work on "Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study" really zeroes in on that specific timing gap we talked about earlier.
Dev: They are essentially confirming that the standard way some MLCs handle data delivery through those I2C transactions introduces measurable delay compared to a direct host-side approach, especially when the system is running under load or contention.
Rosa: The implication here is practical: if you're designing a wearable exoskeleton control system or something safety-critical, you can't just trust the functional result of an embedded AI without rigorously measuring this wire-level latency because it dictates whether the system can physically react in time.
Dev: Exactly, and their findings about how the MLC binary variant isolates that kernel/gpiod latency floor really shows us that when we strip away the ML computation itself, we see exactly what’s slowing down our decision path on this hardware.
Taro: So, for autonomy research, this means if you're relying on sensor fusion where decisions need to be made in milliseconds, you have to account for these I2C protocol overheads because they can unexpectedly inflate the total time it takes for a reaction.
Rosa: And that’s why their finding that the three-transaction I2C read protocol amplifies contention penalties is so important—it tells us exactly what kind of operational stress could push a system over its timing budget.
Dev: It's a very concrete piece of data because they didn't just talk about theory; they used external timestamped measurements to prove that the host pipeline is faster in many situations, which helps us build more realistic performance models for deployment.
Taro: I think the broader impact is forcing the industry to move beyond just reporting ML accuracy and start demanding latency specifications that account for these underlying hardware communication layers when building autonomous agents.
Rosa: It moves the focus from "how good is your classifier?" to "can your entire sensing-to-action loop meet its time constraints?"
Dev: So, in simple terms, the study on this paper confirms that for many scenarios on the Jetson Orin Nano, running inference directly on the host side beats the standard embedded protocol because of communication overhead.
Taro: That really hammers home that when you're building things that interact with a physical world, you need to be meticulous about those communication details, not just the high-level algorithms.
Rosa: It’s about understanding how much time is spent waiting for a signal to travel and be processed by the I2C stack versus the actual intelligence calculation itself.
Dev: Precisely. The paper gives us a concrete measure of that waiting time under stress, which means we can finally start designing systems that are timing-aware from the beginning rather than trying to patch performance issues later on when things go wrong in deployment.
Taro: That seems like a very useful contribution for anyone working on real-time AI where failure modes involve missed deadlines, and it helps us predict those failure modes more accurately.
Rosa: It gives us a solid foundation for performance requirements that aren't just theoretical ideals but are backed by measured wire-level data showing the actual behavior under various operating conditions on this specific hardware setup.
Akul Swami, Dnyaneshwar Sonawane
eess.SY, cs.SY
Submitted: 2026-05-30
Updated: 2026-09-27
Comments: 5 pages, 2 figures, 1 table. v2: revision under review at IEEE Sensors Letters (SENSL-26-06-RL-0906.R1). Code, data, pre-registration: github.com/akulswami/sensor-mlc-latency
Code: https://github.com/akulswami/sensor-mlc-latency
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested
Key concepts
- Wire-Level Latency
- This measures the actual time delay from a physical signal event (like an interrupt edge) to a corresponding electrical change on a wire (like a GPIO pin). It is measured using high-resolution logic analyzers to bypass software timing uncertainties, providing an accurate picture of hardware performance.
- MLC Pipeline
- This refers to the onboard Machine Learning Core running on the sensor itself. The study tested different ways this core operates, specifically comparing a standard version that performs three I2C reads per interrupt against a binary version that skips the read entirely to isolate system overhead.
- I2C Read Protocol Overhead
- The time spent communicating with the sensor via the I2C bus is identified as the main source of delay. The study found that even when comparing different ML processing methods, the extra time required for these three-transaction I2C reads significantly slows down the overall decision process.
- Interrupt-to-Decision Latency
- This is the total time elapsed from when a sensor generates an interrupt signal to when the host system's output GPIO line changes its state based on that interrupt. This metric was used to compare how quickly different processing methods (host versus MLC) can react to incoming sensor data.
Terminology
Summary
Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested conditions. This study uses a Saleae Logic Pro 8 logic analyzer to measure the time from a sensor interrupt edge to the host decision GPIO, demonstrating that the dominant latency contributor is often not the silicon's classification itself but rather I2C read protocol overhead.
Experimental Setup and Platform
The research was conducted on an NVIDIA Jetson Orin Nano Developer Kit running JetPack 6.2.2, utilizing an STMicroelectronics LSM6DSOX 6-axis IMU connected via I2C bus 7 at 400 kHz. The measurement involved instrumenting two specific GPIO lines with a Saleae Logic Pro 8 to capture nanosecond-resolution wire-level timestamps independent of the Jetson clock or SSH jitter. Three distinct pipelines were compared: (a) a host-side decision-tree classifier, (b) the standard MLC bank-switch read protocol involving three I2C transactions, and (c) an MLC binary variant that omits the I2C read entirely.
Pipelines Under Test and Conditions
The study compared these three pipelines under three stress conditions: idle, I2C bus contention, and CPU saturation. The comparison focused on the interrupt-to-decision latency
(sensor INT1 edge to host decision GPIO). Key differences between the pipelines were:
-
Host pipeline parity: Polls the accelerometer at 208 Hz and applies a decision-tree classifier.
-
MLC pipeline (latency test mlc w75): Executes
the three I2C transactions of the bank-switch read
on every INT1 edge, toggling D1 if the output changed. -
MLC binary variant (latency test mlc binary w75): Toggles D1 unconditionally on every INT1 edge without reading MLC0 SRC, isolating the kernel/gpiod latency floor.
Key Findings from Confirmatory Tests
The principal findings established that the host pipeline exhibits lower median wire-level latency than the standard MLC bank-switch pipeline under all conditions.
Specifically:
((
) H1’ (host < MLC at idle): SUPPORTED. The host idle median (321.7 μs) was 359 μs below the MLC bank-switch idle median (681.5 μs), representing a 2.1× faster
host speedup. The shift was calculated as ΔHL = −359.3 μs, with a high statistical significance (p = 1.87 × 10−170).)
) H2’ (host < MLC under I2C contention): SUPPORTED. Under three concurrent I2C hammer processes, the host median rose to 574.5 μs while the MLC rose to 1,325.4 μs, showing a 2.3× advantage
for the host and a shift of ΔHL = −753.2 μs.
) H3’ (MLC degrades more than host): SUPPORTED. The contrast statistic (Δ-of-Δ) was +391.1 μs, indicating that the MLC's three-transaction I2C read protocol amplifies the contention penalty.
) H5’ (CPU stress null for host latency): SUPPORTED. Host median latency increased only slightly from 321.7 μs to 345.0 μs under stress, and a TOST analysis confirmed equivalence within the ±30 μs margin.
) H6’ (CPU stress positive for energy): SUPPORTED. Mean VDD IN rose significantly from 5,206 mW (idle) to 8,626 mW (stress), exceeding the pre-registered +1,000 mW threshold threefold.
Structural Observations and Cadence
Beyond direct latency comparisons, the study characterized structural aspects of the MLC operation. The researchers documented a reproducible 706.5 ms MLC decision cadence,
which bounds full stimulus-to-decision latency. This cadence is consistent with one quarter of the 75-sample × 26 Hz window period.
Furthermore, multimodal distributions were observed at idle for the MLC pipelines, where the p95 distribution reached 1,780.7 μs, suggesting that the upper mode, not the median,
should be used for worst-case latency specifications.
Conclusion on Latency Drivers
The headline finding is that the three-transaction I2C read protocol, not the silicon’s classification,
is the dominant latency contributor to wire-level delay.
Improvements for AI systems
Here are the specific improvements for AI systems based on this research, categorized by capability:
) Improve Real-Time Safety-Critical Decision Systems (e.g., Autonomous Robotics, Wearable Exoskeletons)
-
Enable
True
Low-Latency Control Loops: Implement an architecture where the host system's decision pipeline is prioritized over the on-sensor MLC path when high reliability is required. -
Mitigate Bus Contention Latency: Design communication protocols for IMU data (specifically I2C bank-switched reads) that minimize transaction overhead, as this overhead scales with bus activity rather than classifier complexity.
-
Implement Latency-Aware Scheduling: Integrate the findings on CPU/I2C contention into real-time operating system (RTOS) scheduling to ensure critical decision paths are not starved by background processes or high bus traffic.
-
Predict Worst-Case Latency: Use the characterized 706.5 ms MLC decision cadence as a structural floor for system-level latency budgeting, ensuring that unsynchronized external stimuli do not cause catastrophic timing failures in the overall control loop.
) Enhance Edge Inference Performance and Reliability (for IoT/Industrial Monitoring)
-
Optimize On-Sensor ML Deployment: When on-sensor inference is chosen, utilize the
MLC binary
variant (or equivalent zero-read protocol) to bypass I2C overhead entirely for simple, high-frequency classification tasks where data throughput is prioritized over complex classification depth. -
Refine Stress Testing Metrics: Move beyond simple CPU utilization metrics for edge AI validation; use energy consumption (VDD IN) as a more robust indicator of system stress, as it remains a strong differentiator between benign load and genuine resource saturation.
-
Improve Classifier Stability Under Load: Develop on-sensor ML models that demonstrate superior resilience to I2C bus contention (as the study found classifier reliability is not degraded by contention), allowing for higher confidence in edge deployments where bus traffic is unpredictable.
) Develop Robust Sensor Measurement Verification Frameworks (for Sensor Fusion & Calibration)
-
Establish Pre-Registration Standards: Mandate the use of externally timestamped DOIs (like Zenodo) for all sensor measurement claims to create an audit-defensible chain of custody, ensuring that performance claims are verifiable against a fixed methodology.
-
Distinguish Protocol Overhead from Model Error: When evaluating sensor fusion systems, explicitly separate latency contributions: if performance degrades under high bus load, the system can definitively attribute the degradation to protocol overhead (like I2C read transactions) rather than inherent model instability.
Sources
- Event-Driven On-Sensor Locomotion Mode Recognition Using a Shank-Mounted IMU with Embedded Machine Learning for Exoskeleton Control
- Per-Platform GPIO Overhead in Hardware-Validated Edge ML Inference Timing
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation