Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

arXiv:2606.19998 · cs.RO, cs.AI, cs.CV, cs.LG · Submitted 2026-06-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory".

Dev: Vision-Language-Action (VLA) models are increasingly deployed in robotics, yet they remain black boxes whose physical interactions can cause irreversible harm, necessitating generalizable and interpretable failure detection.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Moving on, let's talk about the paper's title and who the authors are, because understanding the team behind the research often tells us a lot about the direction of this new work.

Dev: I’m curious to see if it’s just a theoretical exercise or if these researchers have actually seen this framework put into practice on physical hardware yet.

Taro: From my perspective, seeing how these information-theoretic concepts map onto concrete robotic behaviors is what matters most for autonomy research.

Rosa: The title itself, "Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory," tells us immediately that the authors are focused on three specific goals: generality, interpretability, and using information theory as the backbone.

Dev: That points toward a system designed to be robust across different setups without needing constant manual tuning or retraining.

Taro: And the fact that they are emphasizing information theory suggests they’re trying to find universal rules governing how these models interact with their environment, which is a big ambition for autonomy research.

Rosa: They achieved something pretty impressive by showing that this framework can achieve eighty-three percent accuracy on real-world tasks where previous detectors were failing entirely.

Dev: That level of cross-domain transfer without retraining is a huge claim, and it speaks to the underlying substrate-independent nature of their metrics.

Taro: If those claims hold up when we look at deployment scenarios, it means we might not need to re-validate every single robot system from scratch for every new application.

Rosa: So, we’re looking at a framework that aims to be a universal diagnostic tool for the entire VLA landscape, which is certainly an ambitious vision.

Dev: It certainly sounds promising for reducing the safety gap mentioned in papers like SafeVLA-Bench, provided the theoretical rigor translates into practical stability.

The paper's summary: Rosa: Now we’re getting into the meat of it: what exactly does this paper propose and how do these information-theoretic concepts translate into a usable system for monitoring VLA control?

Dev: So, in simple terms, they take the VLA control pipeline and model it as a continuous flow of information, and then they derive three specific metrics—action diversity, temporal consistency, and action–state coupling—that capture whether that flow is behaving correctly.

Taro: That’s the core mechanism: they aren't just looking at the final state; they are analyzing the entire trajectory to see *how* the information moves through perception and action.

Rosa: They systematically derive eight potential metrics from various categories—marginal statistics, policy coupling, dynamics, and temporal coherence—but then they narrow those down to these three Tri-Info signals for maximum diagnostic power.

Dev: The paper highlights that these three signals are designed specifically to capture the distinct failure modes we discussed earlier: drift, freeze, and phantom grasp.

Taro: That’s the interpretability part; instead of a vague error code, we get a specific diagnosis like "high action entropy" pointing directly at a drift failure.

Rosa: So, the summary really boils down to creating an information-theoretic dashboard that provides interpretable diagnostics by linking mathematical concepts to observable failures in robot behavior.

Dev: It seems like they’ve successfully formalized the control process as a pipeline, which makes it much easier for us to audit where things are going wrong step by step.

The paper's improvements: Rosa: Let’s look at what the authors claim are the specific improvements in this approach over existing methods, especially when we compare it to other detectors that might rely on simpler, architecture-specific scores.

Dev: The major improvement seems to be moving away from coordinate geometry-based metrics toward metrics that are functionals of the embedding distribution itself, which is what grants them that substrate independence.

Taro: That’s significant because it means the detector isn't tied to a specific neural network architecture, allowing it to transfer across different models and environments without needing retraining.

Rosa: They demonstrated this by showing that Tri-Info reaches eighty-three percent accuracy on real-world tasks where prior detectors just collapsed to chance, which is a strong validation of its generalizability.

Dev: That result really puts it in contrast with embedding-based methods that are architecture–specific, and scoring methods that only manage to relocate the difficulty instead of solving the core issue.

Taro: It shows they’ve found a way to build a detector that addresses the actual physics of failure rather than just looking at superficial performance indicators.

Rosa: And for deployment, they've built an online detection framework featuring per-metric GRU detectors fused together with a late mean-probability fusion.

Dev: The final piece of the puzzle that I like is the use of Functional Conformal Prediction to build a time-varying threshold that accounts for the natural shift in success probabilities during a rollout.

Taro: That dynamic thresholding is clever because it allows the system to flag potential failures significantly earlier than static thresholds, which directly addresses our need for timely intervention.

Conclusion: Rosa: So, to wrap up this discussion on "Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory," we've covered the mechanics of the Tri-Info signals and how they diagnose drift, freeze, and phantom grasp failures.

Dev: It’s clear that by formalizing control as a closed-loop information pipeline gives us a robust way to monitor the system’s behavior without needing constant retraining for new scenarios.

Taro: The paper’s implication is that we can start developing more reliable diagnostic tools for complex AI systems that go beyond just reporting high or low success rates.

Rosa: It delivers interpretable diagnostics by pointing to mode-specific interventions, such as re-injecting exploration or rolling back perception, which gives us actionable steps instead of just a warning.

Dev: It seems like the Tri-Info framework is a simple yet powerful method because it has negligible overhead and still achieves high accuracy even when facing distribution shifts.

Taro: I think the paper’s final message is that we can gain deep, mechanistic understanding of why VLA models fail by analyzing their information flow rather than just observing the output.

Rosa: We’re ready to move on to what this means for our real-world robotic systems and what comes next in this research area.

Jinghan Yang, Yunchao Zhang Wang Yuan Haolun Wang, Jiaming Zhang Zhengyang Hu, Yanchao Yang

InfoBodied AI Lab, The University of Hong Kong

cs.RO, cs.AI, cs.CV, cs.LG

Submitted: 2026-06-18

Updated: 2026-09-30

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Vision-Language-Action (VLA) models are increasingly deployed in robotics, yet they remain black boxes whose physical interactions can cause irreversible harm, necessitating generalizable and

Key concepts

Tri-Info Metrics
These are three specific information measures derived from the system's state and action embeddings. They quantify different aspects of control flow, including how varied actions are taken, how consistent the sequence is over time, and the relationship between states and actions.
Failure Modes
The framework identifies three distinct failure types: drift failures (action entropy surge), freeze failures (entropy collapse), and phantom grasps (drop in state-action mutual information). These modes provide a mechanistic explanation for why VLA models fail during execution.
Substrate-Independent Generalization
TriInfo works by analyzing the distribution of embeddings rather than relying on specific geometric coordinates. This makes the resulting metrics 'substrate-independent,' allowing them to transfer effectively across different VLA architectures and environments without needing to be retrained for each new setup.

Terminology

Summary

Vision-Language-Action (VLA) models are increasingly deployed in robotics, yet they remain black boxes whose physical interactions can cause irreversible harm, necessitating generalizable and interpretable failure detection. This paper introduces TriInfo, a novel information-theoretic framework that formalizes VLA control as a closed-loop information pipeline to derive three complementary signals—action diversity, temporal consistency, and action–state coupling—which capture systematic differences between successful and failed rollouts. TriInfo is shown to be highly generalizable across architectures and environments without retraining, achieving 83% accuracy on real-world tasks where prior detectors fail.

Systematic Information-Theoretic Framework

The authors formalize VLA control as a closed-loop information pipeline, modeling the system as a tuple (S, A, T, p, πw) where state embeddings are extracted via ϕw(st) and action embeddings via ψw(st, τ). They systematically derive eight metrics across four diagnostic categories: marginal statistics (H(St), H(At)), policy coupling (I(St; At)), dynamics (I(At; St+1), I(At, St; St+1), I(St, St+1; At)), and temporal coherence (I(St; St+1), I(At; At+1)). Through correlation analysis, they reduce these eight candidates to three complementary signals: the Triple Information-theoretic (Tri-Info) metrics:

H(At)

[I(At; At+1)]

[I(St, St+1; At)]

Failure Mode Diagnosis and Interpretability

The three Tri-Info signals are designed to capture distinct failure modes, providing a mechanistic account of why failures occur. The paper identifies three representative failure modes illustrated in Figure 1:

  1. A drift failure, characterized by a surge in action entropy (H(At) ↑).

  2. A freeze failure, characterized by the opposite collapse in entropy (H(At) ↓).

  3. A phantom grasp, characterized by a drop in state–action mutual information (MI) (I(St, St+1; At) ↓).

Generalizability and Robustness

TriInfo achieves significant cross-domain transfer capabilities because its metrics are functionals of the embedding distribution rather than coordinate geometry, allowing them to be substrate-independent. The authors demonstrate that TriInfo transfers across VLA architectures, environments, and the sim-to-real gap without retraining, reaching 83% accuracy on real-world tasks where prior detectors collapse to chance. This contrasts with embedding-based methods that are architecture–specific and score-based methods that only relocate the difficulty.

Online Detection Framework

To deploy TriInfo in real time, the authors build a detector featuring two components:

  1. A per-metric GRU detector: For each of the three signals, a single-input GRU models the temporal evolution of failure signatures. The output is calculated as Pξt(Failure xt) = sigmoid(MLP(ht)).

  2. Late mean-probability fusion: The detectors are aggregated by averaging their outputs to produce a final score: P¯ξt = 1/3 Σ Xm∈M Pξt(Failure x(m))t, where M is the set of Tri-Info metrics.

  3. Time-varying threshold via Functional Conformal Prediction (CP): A fixed threshold is inadequate because success scores shift systematically. They use CP to build a dynamic threshold θ(t) that caps how high the score may rise at stage t of a successful trajectory: θ(t) = Quantile1−α n Pξt(Failure xt): ξ ∈ D+calo.

Experimental Validation

The detector was validated across six VLA models and three benchmark environments, including both simulated platforms and two real-world robot tasks (ACT-ALOHA-real). The results show that the Tri-Info detector consistently achieves higher accuracy at both early and final time points compared to baselines. Ablation studies confirm that the three-metric set is optimal, as adding more signals beyond the triplet leads to a degradation in out-of-distribution performance, confirming their complementary nature. The analysis also shows that while temporal modeling saturates detection in-domain, fusion buys robustness and earlier detection under distribution shift.

Conclusion and Limitations

TriInfo successfully formalizes VLA control using information theory to create an interpretable failure mode dashboard. It is a simple yet powerful method that detects failures with strong cross-domain generalization, delivering interpretable diagnostics of the underlying failure modes by pointing to mode-specific interventions such as "re-injecting exploration, rolling back, or re-grounding perception.

Improvements for AI systems

Based on the provided scientific paper, here are the specific improvements that can be made to AI systems by implementing the Tri-Info failure detection framework, along with what those improved systems will be able to do:


The core improvement is moving from black-box deployment monitoring to an interpretable, generalizable failure anticipation system.

  1. Improve Robustness Against Distributional Shift (Generalization):

  2. Improve Early Failure Detection Timeliness (Safety):

  3. Improve Diagnostic Capability for Failure Modes (Interpretability).

Specific System Capabilities:

  1. A VLA model deployed in a novel environment or with a different robot embodiment can be monitored by the Tri-Info detector, which is expected to maintain high accuracy (up to 83% on real-world tasks) where existing detectors collapse. This allows systems to operate safely outside of their original training distribution without needing costly retraining.

  2. The system will detect and diagnose distinct failure modes in real-time:

Narrow the scope of failures from a single failure flag to specific, actionable modes:

  • Detect a Drift Failure (erratic movement/goal abandonment) by monitoring a surge in action entropy, triggering an intervention to re-ground the policy.

  • Detect a Freeze Failure (stalling) by monitoring the collapse in action entropy, potentially triggering a recovery routine or exploratory action injection.

  • Detect a Phantom Grasp (acting as if holding an object it never grasped) by monitoring a drop in state–action mutual information, signaling to the system that the perception-action coupling is broken and demanding re-perception or re-grasping.

  1. The system provides an interpretable dashboard that aligns with visually identifiable dangerous events (as shown in Figure 1). Instead of just reporting Failure Probability, the operator receives a diagnosis like, Warning: High drift signature detected at t=X, which translates directly into a specific intervention strategy (e.g., Re-inject exploration or Roll back to last safe state).

  2. The system achieves early warning by using Functional Conformal Prediction (CP) to dynamically set a threshold that accounts for the natural shift in success probabilities during a rollout. This means the system can flag potential failures significantly earlier than fixed-threshold detectors, providing a crucial window for timely human or autonomous intervention before irreversible harm occurs.

  3. The Tri-Info detector is compatible with deterministic VLA inference and has negligible overhead (under 5ms), meaning it can be integrated directly into the control loop without slowing down the robot's decision-making process.

Sources

Related papers