TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

summary

Video file (mp4)

The gist

Active vision—where a policy controls its own gaze during manipulation—has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the

In short

TAVIS introduces a new evaluation infrastructure to test imitation learning policies that use active vision, where a policy controls its own gaze during manipulation. It compares different active vision tasks and conditions using two task suites and novel metrics like GALT, providing a shared benchmark for this emerging field.

Key concepts

Active Vision
This is when an agent actively moves its viewpoint, such as turning its head, to gather information needed to perform a task. It is different from passive vision where the camera stays still. Active vision allows the robot to control *where* it looks in order to solve problems like finding hidden objects or handling clutter.
GALT (Gaze-Action Lead Time)
This is a new metric that measures how much an agent anticipates its gaze before taking an action. It is calculated as 'thand minus thead,' where 'thead' is when the agent stops looking before grasping, and 'thand' is when it actually completes the grasp. A positive GALT score indicates that the agent was successfully anticipating where it needed to look ahead.
ID/OOD Distribution Splits
These splits are used to test how well a learned policy generalizes. ID (In-Distribution) tests performance on data similar to what it was trained on, while OOD (Out-of-Distribution) tests performance on novel or perturbed scenarios. This helps determine if the policy can reliably extrapolate its skills in new situations.

Terminology used across episodes

This episode discusses

The paper

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning · Read on arXiv

Department of Intelligent Systems · Tilburg University

Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which task types, and under what conditions. We introduce TAVIS, evaluation infrastructure for active-vision imitation learning, with two complementary task suites -- TAVIS-Head (5 tasks, global search via pan/tilt necks) and TAVIS-Hands (3 tasks, local occlusion via wrist cameras) -- on two humanoid torso embodiments (GR1T2, Reachy2), built on IsaacLab. TAVIS provides three evaluation primitives: a paired headcam-vs-fixedcam protocol on identical demonstrations; GALT (Gaze-Action Lead Time), a novel metric grounded in cognitive science and HRI that quantifies anticipatory gaze in learned policies; and procedural ID/OOD splits. Baseline experiments with Diffusion Policy and π 0 reveal that (i) active vision generally helps, consistently across training runs, but benefits are task-conditional rather than uniform; (ii) multi-task policies degrade sharply under controlled distribution shifts on both suites; and (iii) imitation alone yields anticipatory gaze that lands on the task-relevant object, with lead times comparable to those of the human demonstrations, although head motion is less smooth than in the demonstrations, an effect of action chunking that success rate does not reveal. Code and evaluation scripts are released at https://github.com/spiglerg/tavis; demonstrations (LeRobot v3.0; 2200 episodes) and trained baselines at https://huggingface.co/tavis-benchmark.

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning".

Dev: Active vision—where a policy controls its own gaze during manipulation—has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we've got the paper "TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning," which is looking at how a policy can control its own gaze during manipulation. This means it learns to look where it needs to look while physically moving, which is a big step for imitation learning.

Dev: And the authors, Giacomo Spigler from Tilburg University Netherlands, they've built this infrastructure specifically because there wasn't a shared way to compare different approaches or quantify exactly how much active vision contributes across different tasks. That lack of comparison is what TAVIS addresses directly.

Taro: It’s interesting that they’re focusing on active vision as a key capability for imitation learning, because it seems like something that really matters when a robot has to interact with the real world dynamically instead of just following pre-programmed paths.

Rosa: Exactly, and looking at the title again, it sets up this evaluation infrastructure to see what active vision helps with and under what conditions. It’s not just about whether a policy can pick an object; it's about *how* it looks for that object while moving.

Dev: And they've structured their benchmark around two specific task suites, TAVIS-Head for global search using head reorientation, and TAVIS-Hands which focuses on local occlusion handled by wrist cameras. That separation seems really smart for isolating different kinds of active vision needs.

Taro: I’m thinking about the implications of having these distinct suites; it suggests that active vision isn't one-size-fits-all and we need to test it in contexts where you're searching broadly versus when you're just dealing with something right in front of your hand.

Rosa: Right, and they are using two humanoid torso embodiments, GR1T2 and Reachy2, both operating under a unified nineteen-dimensional canonical action space. That’s important because it lets us compare how different robots handle the same active vision tasks without getting bogged down by the robot's specific hardware differences.

Dev: The authors also provide three evaluation primitives: a paired headcam-vs-fixedcam protocol, GALT, and ID/OOD distribution splits. The GALT metric is particularly interesting because it’s a kinematic measure grounded in cognitive science that tries to quantify anticipatory gaze.

Taro: Quantifying that anticipation with GALT—defining it as the difference between the time of final pre-grasp fixation and grasp completion—that sounds like a solid way to measure if the policy is actually looking ahead, rather than just reacting.

Rosa: And they also use ID and OOD distribution splits to distinguish between interpolation, where the robot does what it was trained on, versus extrapolation when it encounters something new or outside its expected range. This helps us understand generalization issues better.

Dev: Baseline experiments with Diffusion Policy and π0 showed that active vision generally provides benefits, but those benefits are task-conditional rather than universal across every scenario tested in the TAVIS benchmark for Active Vision and Anticipatory Gaze in Imitation Learning <ref:2605.07943#pg2>.

Title and authors: Taro: That finding about task conditionality is important because it suggests that we can't just assume active vision helps everywhere; we have to know which manipulation challenges benefit from it most significantly, like on conditional-pick tasks where the boost is noted as plus twenty-eight percentage points on GR1T2 and plus forty-five percentage points on Reachy2.

Rosa: That leads us into the idea of enhancing manipulation robustness by training policies specifically under those task conditions, which sounds like a clear path for improvement.

Dev: But the paper also flagged that multi-task policies tend to degrade sharply when subjected to controlled distribution shifts on both suites, which is a cautionary note for deployment. For instance, the mean success rate for TAVIS-Head drops from forty-three point zero percentage points in in-distribution scenarios down to twenty-five point zero percentage points when tested under an out-of-distribution spatial shift.

Taro: That sharp degradation under OOD shifts tells us that while active vision might help in controlled settings, the system isn't inherently robust when things deviate significantly from what it trained on, which is a big hurdle for real-world autonomy.

Rosa: And then they found something interesting about imitation alone: the policies acquired anticipatory gaze just from imitation, and the median lead times were comparable to those of the human teleoperator reference within about one hundred eighty milliseconds on a scale of two point one to two point seven seconds for TAVIS-Head tasks.

Dev: That comparison with the human teleoperator reference via GALT distributions matching within that timeframe suggests that imitation alone is already teaching some level of anticipatory behavior, which makes the contribution of active vision very specific to certain manipulation challenges.

Taro: So, if imitation gives you a baseline for anticipation, then active vision seems to fine-tune or enhance that anticipation based on the task demands rather than just being an additive feature.

Rosa: That’s a good way to frame it; we're moving from simple imitation toward policies that actively control their gaze during manipulation based on the specific environment they encounter. This leads us into what the authors suggest for improvement in this work.

Dev: The suggested improvements seem centered around integrating active gaze control directly into imitation learning policies and enhancing manipulation robustness by training under task-conditional scenarios, which addresses that finding about variable benefits.

Taro: I agree with focusing on task-conditional training; if we know when the robot needs to search broadly versus when it needs local occlusion handling, we can tailor the training to maximize performance in those specific roles.

Rosa: And there’s also the idea of improving generalization under distribution shifts, which tackles that sharp drop we saw in the results when moving from in-distribution to out-of-distribution data for both suites.

Title and authors: Dev: Furthermore, they propose using GALT as a way to make robot behavior more legible and communicative by having metrics that quantify anticipatory gaze, which can give us better insights into *why* the AI is looking where it is.

Taro: That metric idea for GALT seems really useful because it moves beyond just success or failure and gives us a kinematic measure of the cognitive process—the anticipation itself.

Rosa: And I think making robot behavior more legible through these metrics has broad implications for human-robot interaction, as we can better understand the AI's internal planning process.

Dev: From an engineering side, having these clear evaluation primitives like GALT and distribution splits gives us concrete ways to test and debug latency and failure modes in a way that goes beyond just looking at final grasp success rates.

Taro: Thinking about the future, if we can get active vision policies to generalize better under distribution shifts, it opens up possibilities for robots operating in less controlled environments where the visual scene changes constantly.

Rosa: Absolutely, and I wonder how long these policies could actually stay reliable outside of a perfectly controlled lab setting before that degradation sets in? That’s a practical question we have to ask.

Dev: And from an engineering standpoint, we need to keep monitoring the loop rate and latency when these active gaze mechanisms are running; if the control loop is too slow, all that anticipation becomes useless and potentially dangerous.

Taro: The paper suggests that understanding these limitations is key, so we can build systems that are more resilient when the world misbehaves because we’ve identified exactly where the current methods fail.

Rosa: So, to wrap up this discussion on "TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning," it provides a solid framework for systematically evaluating active vision capabilities across different contexts using task suites like TAVIS-Head and TAVIS-Hands.

Dev: We see that the main takeaway is that active vision benefits are highly dependent on the specific manipulation task, and we need better ways to handle distribution shifts when moving beyond the training data.

Taro: And I think it points toward a future where imitation learning systems aren't just mimicking actions, but truly learning to control their own perceptual focus in anticipation of future needs.

Rosa: It’s exciting stuff because it moves us closer to creating robots that can handle much more complex, dynamic interactions in the physical world. We'll keep an eye on how these policies perform when they get out into the field, though.

Dev: And we need to focus on making sure the underlying systems running these policies have low latency so that all this active gaze control translates into smooth, reliable physical movement.

Taro: Indeed, understanding those limitations and building systems that can handle unexpected situations is what will really determine how far this technology gets in practical applications.

The paper's summary: Rosa: So, TAVIS is essentially building this whole setup to systematically test how active vision—that’s when a policy actively controls where it looks during manipulation—actually helps imitation learning policies perform under different circumstances.

Dev: Exactly, Rosa; the core idea is that there isn't really a standard yardstick yet for measuring this specific capability, so they put together two task suites and some unique metrics to compare how these active vision systems behave.

Taro: I’m really interested in what this means for real-world autonomy because it suggests that the benefit of looking around is totally dependent on the job, like it isn't always helpful whether you're searching globally or just dealing with a nearby object.

Rosa: That's right; they set up TAVIS-Head to test those big, global searches using head movement, and TAVIS-Hands for when you need that extra local peek capability because the wrist camera is needed.

Dev: And they introduce GALT, which is this kinematic measure where they look at how much time a policy anticipates its next move based on its gaze—basically quantifying that anticipatory gaze we talked about earlier.

Taro: That GALT metric sounds like it gives us a real window into the robot's internal planning process; it moves beyond just seeing if the final grasp worked, and tells us *how* it was looking ahead.

Rosa: And they also have these ID and OOD distribution splits, which lets them see how well those policies actually generalize when things get slightly different or unexpected compared to what they saw during training.

Dev: That's crucial because we saw some sharp drops in performance when the environment shifted from the training data, so seeing that quantified through those splits helps us understand where the failure modes are happening.

Taro: The finding that imitation alone already produces some anticipatory gaze with lead times similar to human teleoperators is really interesting; it suggests the baseline isn't zero for this kind of behavior.

Rosa: It means active vision isn't just adding a feature; it seems to be fine-tuning or enhancing an existing capacity based on the specific demands of the task at hand.

Dev: So, while imitation gives you a starting point, TAVIS helps us pinpoint exactly when and how much active vision adds real value across different manipulation scenarios.

Taro: If we can use these metrics to understand that task conditionality better, we can start training policies that are robust not just to the expected tasks but also to those tricky situations where things go sideways in the real world.

Rosa: And if this infrastructure helps us build policies that are more aware of their perceptual focus in anticipation of future needs, it opens up a lot of avenues for creating robots that handle much more dynamic and complex physical interactions.

The paper's improvements: Taro: So, the paper lays out some clear directions for how we can actually make these active vision policies better than what they currently are by focusing on those task-specific training methods and distribution shifts we discussed earlier.

Rosa: Exactly; they suggest integrating active gaze control more deeply into the imitation learning process itself so that it’s not just an add-on but part of how the policy learns to navigate its visual search space.

Dev: I agree with Rosa; focusing on those task-conditional scenarios means we stop training a single policy and instead create a set of specialized policies, which should drastically improve performance in those specific manipulation roles.

Taro: And they also emphasize making generalization under distribution shifts more robust, which is critical because we know that sharp drop in performance when things get out of spec is a major issue for real-world deployment.

Rosa: That makes sense; if the policies can handle those sudden environmental changes better, we can start thinking about them operating in less controlled settings for longer periods without needing constant retraining.

Dev: From an engineering standpoint, this suggests that instead of just brute-force training on the whole dataset, we should use these distribution splits to identify exactly which types of visual perturbations are most damaging to the system's control loop.

Taro: And they also propose using GALT more actively as a way to make robot behavior more legible, which I think is huge because it gives us a quantifiable measure of the cognitive intent behind the movement.

Rosa: Legibility through metrics is powerful; if we can see *why* an AI looked where it did, we gain much better insight into its planning than just looking at the final success rate.

Dev: That’s a good point, Rosa; and for us in engineering, those quantitative insights from GALT can help us debug latency issues related to gaze control loops because we have a metric tied directly to that anticipation.

Taro: So, by combining task-specific training with better generalization metrics and legibility tools like GALT, we’re moving toward systems that are not just successful in the lab but are genuinely resilient when the world throws curveballs.

Rosa: It seems like they’re pushing for a system where the AI learns to control its own perceptual focus based on what it needs to do next, rather than just reacting blindly.

Dev: That level of active, goal-driven gaze control is something we need to see implemented with very low latency if we're going to put these systems into any kind of operational setting where timing matters.

Taro: If these improvements lead to more robust and understandable AI, the impact on how we design autonomous agents for complex physical tasks could be significant.

Rosa: It’s exciting stuff because it moves us closer to creating robots that can handle much more dynamic interactions in the physical world, provided those policies stay reliable outside of a perfectly controlled lab setting for a reasonable amount of time.

Conclusion: Rosa: So we've covered how TAVIS sets up this comprehensive evaluation infrastructure for egocentric active vision policies by comparing different task types and conditions using things like GALT and distribution splits.

Dev: Right, Rosa; we established that these benchmarks are really pushing us to look at active vision as a measurable capability rather than just an assumed feature in imitation learning.

Taro: It’s fascinating how they’ve managed to isolate the effect of active vision on things like clutter disambiguation versus local occlusion, which tells us a lot about the underlying cognitive demands of manipulation.

Rosa: And I think this work is really important because it gives us a structured way to test these complex visual skills, moving beyond simple success rates to look at how policies anticipate their needs.

Dev: Absolutely; and from an engineering standpoint, having those concrete metrics like GALT and the paired headcam setups means we can actually diagnose latency issues related to gaze control in a way that’s currently very hard.

Taro: I agree; if we can get these systems to handle distribution shifts better, it opens up possibilities for robots operating in less controlled environments where the visual scene changes constantly, which is where autonomy really needs to be tested.

Rosa: Ultimately, TAVIS gives us a solid framework for systematically evaluating active vision capabilities across different contexts and helps us understand *when* and *why* that active gaze control is actually beneficial.

Dev: So, while the paper lays out some clear directions for how we can make these active vision policies better than what they currently are by focusing on those task-specific training methods and distribution shifts we discussed earlier, it's a strong foundation for future work.

Taro: I think this research points toward a future where imitation learning systems aren't just mimicking actions, but truly learning to control their own perceptual focus in anticipation of future needs, which is where the real autonomy lies.

Rosa: It’s exciting stuff because it moves us closer to creating robots that can handle much more dynamic interactions in the physical world, provided those policies stay reliable outside of a perfectly controlled lab setting for a reasonable amount of time.

Dev: And we need to keep monitoring the loop rate and latency when these active gaze mechanisms are running; if the control loop is too slow, all that anticipation becomes useless and potentially dangerous for real-time operation.

Taro: Indeed, understanding those limitations and building systems that can handle unexpected situations is what will really determine how far this technology gets in practical applications.

More episodes

← Home