TRIAGE: Type-Routed Interventions via Aleatoric-Epistemic Gated Estimation in Robotic Manipulation and Adaptive Perception -- Don't Treat All Uncertainty the Same

arXiv:2603.08128 · cs.RO, cs.LG · Submitted 2026-03-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "TRIAGE: Type-Routed Interventions via Aleatoric-Epistemic Gated Estimation in Robotic Manipulation and Adaptive Perception -- Don't Treat All Uncertainty the Same".

Dev: Most uncertainty-aware robotic systems collapse prediction uncertainty into a single scalar score and use it to trigger uniform corrective responses,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: , welcome everyone to the show today; we’re talking about a paper called "TRIAGE: Type-Routed Interventions via Aleatoric-Epistemic Gated Estimation in Robotic Manipulation and Adaptive Perception -- Don't Treat All Uncertainty the Same." Rosa, you open this up for us—what’s the core idea here, and why is this decomposition of uncertainty important?

Rosa: Well, the main thesis of TRIAGE is that most existing uncertainty-aware robotic systems just lump all their prediction uncertainty into one single score and react uniformly to it. This paper argues that’s a mistake because it hides whether a system is struggling because its observations are noisy or because its internal model doesn't match the real world dynamics. They introduce a framework that breaks this down into two distinct signals: aleatoric uncertainty, which relates to sensor noise, and epistemic uncertainty, which points to mismatches in the learned model or dynamics.

Dev: That distinction is exactly what I'm interested in from an engineering standpoint; treating them separately opens up different ways to handle failures. So, if I understand correctly, they’re proposing a post hoc framework that uses these separate signals to regulate the system's response at inference time instead of just using one aggregate score?

Taro: Exactly; it moves beyond just reporting uncertainty and actually uses that information to decide what action to take. The paper claims this decomposed approach improves manipulation robustness significantly, showing a jump from sixty-three point eight percent up to ninety-four point two percent when compared to monolithic methods <ref:2603.08128#pg0>.

Rosa: That jump in robustness sounds substantial; I'm curious about the practical application outside of a controlled lab setting; could this framework actually perform well when the robot is dealing with unpredictable, real-world environments for extended periods?

Dev: That’s a big question, Rosa; we need to look at the latency and loop rate implications here. If we’re introducing two separate estimation processes—one for observation noise and one for dynamics mismatch—how does that affect the overall system performance under tight timing constraints?

Taro: The structure of the framework is designed to be lightweight post hoc decomposition, which suggests it’s not adding a massive computational burden during operation, which is good when you consider resource-constrained settings <ref:2603.08128#pg2>. They even showed a reduction in tracking compute by fifty-eight point two percent on MOT17 without losing much accuracy <ref:2603.08128#pg0>.

Rosa: A compute reduction that keeps the detection quality within zero point four percent is impressive; it suggests this isn't just theoretical work confined to simulations, but something that can be deployed where processing power is limited.

Paper summary: Dev: And from a control engineering view, the idea of type-specific interventions—using observation recovery when aleatoric uncertainty spikes and action dampening when epistemic uncertainty rises—that sounds like a very targeted way to manage failures in real-time. I wonder about the failure modes if one of those two signals is misestimated?

Taro: The paper addresses that by showing the resulting signals are nearly orthogonal, with an empirical correlation of only zero point zero four eight <ref:2603.08128#pg0>, which confirms they capture distinct disturbance mechanisms, meaning the system isn't relying on a single signal to guide its decisions.

Rosa: That orthogonality is key; it means you can have different corrective actions triggered by the two signals without them interfering with each other in an unwanted way. So, for listeners who might be interested in autonomous systems, what does this mean when the world misbehaves unexpectedly?

Dev: When the world misbehaves, this system doesn't just freeze or guess; it attempts to figure out *why* things are going wrong—is it because the sensor is sending bad data, or is the physical system behaving differently than expected? This allows for much smarter adaptation than a simple threshold-based response.

Taro: Precisely, and this principle extends beyond robotics to other areas, like language models where they apply conformal abstention policies to improve risk management under distribution shift <ref:2603.08128#pg1>. The concept of type-specific interventions is a general principle for handling complex uncertainty across different domains.

Rosa: It sounds like the title, "TRIAGE: Type-Routed Interventions via Aleatoric-Epistemic Gated Estimation in Robotic Manipulation and Adaptive Perception -- Don't Treat All Uncertainty the Same," really captures the essence of moving away from that single scalar approach.

Dev: I think the authors are focused on showing that separating these signals allows for a much more nuanced control loop, which is something we need when dealing with high-speed, real-time systems where latency matters immensely <ref:2603.08128#pg1>.

Taro: And their method of calibrating the aleatoric score using Mahalanobis distance in observation space and epistemic uncertainty using a noise-robust dynamics ensemble trained on clean and noise-augmented transitions is a solid way to ground these signals in measurable physical reality <ref:2603.08128#pg1>.

Rosa: So, we’ve talked about the core claim—that decomposition matters—and how the framework handles different types of disturbances. Now, let's look at what this means for the broader impact of TRIAGE and where it goes next.

Dev: I think one big implication is that resource-constrained systems can gain significant autonomy by making smarter decisions about when to adapt their models or when to reduce control effort based on the specific type of uncertainty they are facing <ref:2603.08128#pg2>.

Taro: The potential impact is in building more resilient agents; if a robot encounters a sudden change in friction, it knows that's an epistemic signal, so it will dampen its control actions rather than blindly trying to correct the movement based on bad sensor readings. This level of situational awareness is what we need for truly autonomous systems.

Paper summary: Rosa: I’m really excited by the idea that this structure allows for a fifty-eight point two percent compute reduction in tracking inference while preserving accuracy, which points toward real-world deployment viability rather than just academic curiosity <ref:2603.08128#pg0>.

Dev: Speaking of deployment, the paper does mention a specific limitation; they state that the two estimators require a reference distribution representing nominal system behavior, achieved through calibration via nominal rollouts, and they have to perform a short nominal rollout of three hundred steps to set their epistemic threshold <ref:2603.08128#pg1>. That reliance on that pre-calibration step could be a challenge in environments where the robot's initial state is highly uncertain or constantly shifting in an unknown way.

Taro: That calibration requirement is a fair point; it means the system needs some initial time to learn what "nominal" looks like before it can reliably distinguish between sensor noise and true dynamics shifts, which isn't always guaranteed in chaotic, unpredictable real-world scenarios.

Rosa: So while the structure is powerful for handling known types of uncertainty, we have to be mindful that its performance hinges on getting that initial reference distribution right, which is a hurdle for widespread adoption in unstructured settings.

Dev: From a loop rate perspective, the paper shows how these signals guide actions differently under pure sensor perturbation versus pure dynamics shift; this implies that the system can handle different failure modes at different operational speeds if we tune those thresholds correctly <ref:2603.08128#pg0>.

Taro: The long-term implication is that we might see a trend where uncertainty management moves from being a black box aggregation to something structured and interpretable, allowing us to debug system failures much more effectively, whether in manipulation or in perception systems <ref:2603.08128#pg1>.

Rosa: So, to wrap up this discussion on TRIAGE: we’ve seen how decomposing uncertainty into aleatoric and epistemic components provides a principled way for robots to react differently to noise versus model mismatch.

Dev: And the immediate promise is that this leads to more robust performance across compound perturbations while simultaneously cutting down the computational load during tracking inference <ref:2603.08128#pg0>.

Taro: The real world implication is that we gain a tool for creating agents that don't just react, but understand the nature of the disturbance they are experiencing, which opens up new avenues for complex agentic reasoning systems <ref:2603.08128#pg1>.

Rosa: It’s certainly a framework that tackles uncertainty from a structural standpoint rather than treating it as just a single number to be minimized or maximized.

Conclusion: Rosa: I think the title itself is really descriptive because it immediately tells you the core mechanism: routing interventions based on whether the system is dealing with observation noise or dynamics mismatch. The authors clearly want to emphasize that treating all uncertainty as one blob isn't sufficient for robust operation in these kinds of tasks.

Dev: I agree, Rosa; from an engineering standpoint, that separation is crucial because it dictates *what* you actually do next, which is what I care about most in terms of control loop stability and latency. The authors are proposing a way to make the system's response conditional on the source of the uncertainty.

Taro: And for autonomy research, this means if a robot encounters a situation where its sensors are just providing noisy readings, it might focus on observation recovery, whereas if the underlying physical model is simply wrong about how things move, it should focus on action dampening. That tailored approach to failure handling is what makes it interesting for real-world unpredictable environments.

Rosa: Exactly; that tailored handling suggests a much more intelligent way for robots to recover from errors than just a general uncertainty score would allow. It moves the system from a reactive stance to one that understands the nature of the problem occurring at any given moment.

Dev: That distinction between observation noise and model mismatch directly impacts how we design our control loops; knowing which signal is high lets us decide whether to filter input or reduce actuator effort, which has huge implications for real-time performance.

Taro: It opens up a new way to think about agentic reasoning where the system can intelligently decide whether it needs better data from the world or a more accurate internal plan before taking action.

Rosa: So, in simple terms, TRIAGE is giving robots two distinct ways to react when they feel uncertain: one reaction for bad sensing and another for bad understanding of physics.

Dev: That’s the simplest way to put it; it moves away from a single metric toward actionable intelligence based on the root cause of the error.

Taro: And that structural approach, separating these two types of disturbances, is what really makes this work powerful for complex scenarios where things can go wrong in multiple ways at once.

Rosa: It’s certainly a framework that tackles uncertainty from a structural standpoint rather than treating it as just a single number to be minimized or maximized. This idea of targeted intervention is something we need to explore further when thinking about how these systems will operate long-term outside of controlled settings.

University of Illinois at Chicago 2 · Intel Labs

cs.RO, cs.LG

Submitted: 2026-03-09

Updated: 2026-03-09

Journal ref: 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: Most uncertainty-aware robotic systems collapse prediction uncertainty into a single scalar score and use it to trigger uniform corrective responses, which obscures whether uncertainty arises from

Key concepts

Aleatoric Uncertainty
This uncertainty measures noise inherent in the sensor readings themselves, representing deviations from a clean observation distribution. When high, it signals that the measurement is likely corrupted, prompting actions like re-sampling the sensor model to get a cleaner view of the state.
Epistemic Uncertainty
This uncertainty captures errors arising from mismatches between the learned system model and the true physical dynamics. It is detected using an ensemble method that isolates model errors from measurement noise, triggering actions like reducing control strength to compensate for inaccurate predictions.
Type-Specific Interventions
The framework uses the distinct signals to trigger different corrective behaviors based on their source. High aleatoric uncertainty triggers observation recovery, while high epistemic uncertainty triggers action dampening. This separation ensures that the system responds appropriately to whether the problem is a sensor issue or a dynamics issue.
Mahalanobis Density Model
This model is used to estimate aleatoric uncertainty by measuring how far an observed state deviates from the expected nominal state distribution. A large distance indicates that the current observation is statistically unlikely under normal conditions, serving as the primary signal for observation recovery.

Terminology

Summary

Most uncertainty-aware robotic systems collapse prediction uncertainty into a single scalar score and use it to trigger uniform corrective responses, which obscures whether uncertainty arises from corrupted observations or from mismatch between the learned model and the true system dynamics. This work introduces a lightweight post hoc framework that decomposes uncertainty into aleatoric and epistemic components and uses these signals to regulate system responses at inference time.

The gist

The decomposed controller improves task success from 59.4% to 80.4% under compound perturbations in robotic manipulation and reduces tracking compute by 58.2% on MOT17 while preserving detection quality within 0.4%.

How it works: Uncertainty Decomposition and Estimation

The framework separates uncertainty into two distinct signals that correspond to different disturbance mechanisms: aleatoric uncertainty captures deviations in the observation distribution, while epistemic uncertainty captures dynamics mismatch. Aleatoric uncertainty is estimated from deviations in the observation distribution using a Mahalanobis density model, which measures deviation from the nominal state distribution. This signal guides observation recovery when high, triggering actions such as resampling the sensor model to produce a cleaner observation.

How it works: Epistemic Uncertainty and Dynamics Mismatch

Epistemic uncertainty is detected using a noise robust forward dynamics ensemble that isolates model mismatch from measurement corruption. This ensemble is trained on clean and noise-augmented transitions with clean targets, allowing dynamics mismatch to be distinguished from sensor corruption. The epistemic signal triggers action dampening when high, which involves reducing the control magnitude to mitigate the impact of model mismatch while preserving the policy structure.

How it works: Type-Specific Interventions and Orthogonality

The two signals are designed to be functionally distinct, aligning with the separation between sensor perturbations (affecting observations) and dynamics shifts (affecting transitions). The paper demonstrates that the resulting signals exhibit low empirical correlation (r = 0.048), confirming that the two uncertainties capture distinct disturbance mechanisms. This near-orthogonality enables type-specific responses: High aleatoric uncertainty triggers observation recovery, while high epistemic uncertainty moderates control actions.

How it works: Application in Robotic Manipulation and Perception

The decomposed controller operates on different components of the closed loop based on the signals. Under sensor perturbation only, Eq. (8) activates (observation recovery), whereas under dynamics shift only, Eq. (9) activates (action dampening). Under compound perturbations, both interventions may activate concurrently, addressing each disturbance source independently. In adaptive perception for tracking on MOT17, the signals guide adaptive model selection, where epistemic spikes trigger model scaling to larger detectors during representation shifts while ignoring aleatoric elevation caused by measurement noise.

How it works: Calibration and Runtime Adaptation

Both estimators require a reference distribution representing nominal system behavior, achieved through Calibration via nominal rollouts. The aleatoric score is computed as the Mahalanobis distance, and the epistemic signal uses prediction error during closed-loop execution compared to an ensemble mean prediction. To compensate for temporal correlations causing large errors in closed-loop execution, a short nominal rollout of Tcal=300 steps is performed to set the epistemic threshold as τepis = Q95[σ(t)epis]Tcal t=1 × me, ensuring runtime calibration maintains low trigger rates under nominal operation while preserving sensitivity to dynamics shifts.

How it works: Ablation and Performance Gains

Experiments validate the framework's superiority over monolithic methods. The "Total-U baseline aggregates the two uncertainty signals using (σalea > τalea) ∨ (σepis > τepis), which is shown to produce mismatched responses" under pure dynamics shift. In contrast, the Decomposed controller achieves superior performance, with a gain of up to 21.0 pp under compound perturbations compared to the Total-U baseline. Furthermore, the framework avoids cross-contamination that limits its compound performance, maintaining stability across varying dampening strengths and perturbation types where other methods degrade monotonically. In tracking inference, this decomposition leads to a 58.2% compute reduction relative to always using YOLOv8-XL while preserving tracking accuracy.

How it works: Structural Limitations of Monolithic Uncertainty

The paper explicitly shows the structural limitations of monolithic uncertainty, demonstrating that collapsing uncertainty into a single scalar can reduce performance below the vanilla controller under dynamics shift. The decomposed approach proves that separating uncertainty sources improves both corrective control and inference-time resource allocation, confirming that preserving the source of uncertainty enables responses that remain aligned with the underlying disturbance. This principle extends to other domains, such as trajectory-level risk assessment in agentic reasoning systems.

Conclusion

The work presents a post hoc framework that decomposes uncertainty into aleatoric and epistemic components and uses these signals to guide system responses.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for existing AI systems and what those improved systems can achieve:


) Improving Robotic Manipulation Systems (Control Loop):

The core improvement is replacing monolithic uncertainty metrics with a structured decomposition into aleatoric and epistemic signals to govern type-specific corrective actions.

  1. The system can now perform Type-Specific Interventions:

  2. When the observation distribution deviates from the nominal (high aleatoric uncertainty), the system automatically triggers an observation recovery mechanism (resampling from a physics state) to correct sensor noise without altering the control input.

  3. When dynamics mismatch is detected (high epistemic uncertainty), the system applies action dampening to moderate control outputs, preventing actions inconsistent with true system dynamics shifts, without requiring policy retraining.

  4. The improved controller achieves significantly better performance under compound perturbations (up to +21% improvement over the Total-U baseline) by addressing sensor corruption and dynamics mismatch independently, whereas aggregated methods degrade under compound stress.

) Improving Adaptive Perception/Tracking Systems:

The system can now perform Uncertainty-Guided Model Selection during inference at runtime.

  1. The perception module dynamically selects the appropriate model capacity (from Nano to XLarge YOLOv8 detectors) based on the uncertainty signals derived from observation noise versus representation mismatch.

  2. When aleatoric uncertainty is high (indicating measurement noise/blur), the system avoids unnecessary computational load by sticking to a smaller, faster model.

  3. When epistemic uncertainty spikes (indicating viewpoint shifts, occlusion, or abrupt illumination changes), the system escalates to a larger model capacity to capture necessary representation details.

  4. This adaptive selection reduces average inference compute by 58% on MOT17 while maintaining detection quality within 0.4% of the largest backbone, optimizing resource allocation specifically when representation mismatch is present, rather than just when overall uncertainty is high.

Abstract

Most uncertainty-aware robotic systems collapse prediction uncertainty into a single scalar score and use it to trigger uniform corrective responses. This aggregation obscures whether uncertainty arises from corrupted observations or from mismatch between the learned model and the true system dynamics. As a result, corrective actions may be applied to the wrong component of the closed loop, degrading performance relative to leaving the policy unchanged. We introduce a lightweight post hoc framework that decomposes uncertainty into aleatoric and epistemic components and uses these signals to regulate system responses at inference time. Aleatoric uncertainty is estimated from deviations in the observation distribution using a Mahalanobis density model, while epistemic uncertainty is detected using a noise robust forward dynamics ensemble that isolates model mismatch from measurement corruption. The two signals remain empirically near orthogonal during closed loop execution and enable type specific responses. High aleatoric uncertainty triggers observation recovery, while high epistemic uncertainty moderates control actions. The same signals also regulate adaptive perception by guiding model capacity selection during tracking inference. Experiments demonstrate consistent improvements across both control and perception tasks. In robotic manipulation, the decomposed controller improves task success from 59.4% to 80.4% under compound perturbations and outperforms a combined uncertainty baseline by up to 21.0%. In adaptive tracking inference on MOT17, uncertainty-guided model selection reduces average compute by 58.2% relative to a fixed high capacity detector while preserving detection quality within 0.4%. Code and demo videos are available at https://divake.github.io/uncertainty-decomposition/.

Sources

Related papers