Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks
summary
The gist
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks, but this work proposes a benchmark-agnostic evaluation framework to measure behavioral
In short
This work evaluated how input perturbations affect Vision-Language-Action (VLA) models during robot manipulation tasks, moving beyond simple task success rates. The study found that successful trajectories are not robust to these changes; perturbations alter the execution behavior of successful paths. This means task success alone does not capture how a robot actually performs a successful action.
Key concepts
- Task Success Rate (TSR)
- A measure of whether a robot successfully completes its assigned manipulation task. It's the standard way to judge if an AI model achieved its goal, but this paper argues it is insufficient because it doesn't reveal *how* the success was achieved.
- Behavioral Robustness
- The ability of a VLA model to maintain consistent and similar execution metrics (like movement smoothness or trajectory path) even when the input data is slightly altered by noise or perturbations. This measures how stable the robot's actual movements are under stress.
- Trajectory-level Metrics
- Specific quantitative measurements taken from a successful robot movement, such as Cartesian jerk (how fast the end-effector accelerates), total gripper movement, and path length. These metrics describe the physical quality and style of the successful execution.
- Behavioral Variability (MAD)
- A statistical measure used to quantify how much a set of successful movements differs from one another under a specific condition. High variability suggests that while success is achieved, the robot's execution style is inconsistent.
Terminology used across episodes
This episode discusses
- Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks · Paper Radio
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Colosseum V2: Benchmarking Generalization for Vision-Language-Action Models · Paper Radio
- RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- OpenVLA: An Open-Source Vision-Language-Action Model
- MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
- vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models · Paper Radio
- Do World Action Models Generalize Better than VLAs? A Robustness Study
The paper
Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks · Read on arXiv
Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia
University of Edinburgh · Politecnico di Milano
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks".
Dev: Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're discussing the title and authors of "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," and what that really implies for our work.
Dev: The authors are Higham, Izzo, Matteucci, and Suglia, and the paper essentially asks if just achieving a task success rate is sufficient when we consider real-world disturbances.
Taro: I think it's important because they’re pushing the idea that performance in the lab doesn't automatically translate to reliable behavior when environmental conditions shift unexpectedly.
Rosa: That’s right, Taro, and they are proposing this benchmark-agnostic evaluation framework to characterize how successful trajectories behave under perturbation.
Dev: It seems like their main point is that we need metrics that capture the nature of the execution—like how smooth or fast it was—instead of just a single success score.
The paper's summary: Rosa: So, to summarize what they’ve done in "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," they took existing benchmarks like LIBERO and LIBERO-Plus and extended them to test three state-of-the-art VLA models.
Dev: They evaluate these models across four task suites and seven different perturbation conditions, looking at both the typical successful behavior and how variable that behavior is.
Taro: They are calculating metrics like duration, total gripper movement, Cartesian jerk, and joint jerk to get a richer picture of the robot's execution style.
Rosa: They found that perturbations can definitely change the behavior of those successful trajectories, which is something TSR doesn't capture on its own.
Dev: Specifically, they calculated "task-level percentage change in the mean metric value" relative to a no-perturbation baseline for all these metrics.
The paper's improvements: Rosa: Now, regarding the improvements they suggest for this work, it seems they are moving away from relying only on TSR and instead proposing a suite of behavioral metrics that characterize motion smoothness, efficiency, and gripper behavior.
Dev: They propose calculating things like the duration in seconds to isolate task completion time from inference latency or measuring total gripper movement in meters to capture all the open and close motion.
Taro: I see them focusing on jerkiness, both Cartesian and joint jerk, which should give us a good measure of trajectory smoothness under stress.
Rosa: And they also look at path lengths—Cartesian path length for how far the end-effectors move in physical space, and joint path length for how much the arm configuration changes over time.
Dev: They quantify this by calculating both mean and P95 versions of those jerk metrics to see not just the average behavior but also the upper tail of what happens during execution.
Conclusion: Rosa: So, to wrap things up on "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," they conclude that successful trajectories are not behaviorally invariant to perturbations, showing that the model-dependent sensitivity is quite pronounced.
Dev: They also found correlations between TSR degradation and various behavioral metrics, like total gripper movement and mean Cartesian jerk, suggesting success isn't fully capturing execution quality.
Taro: And interestingly, they noted that the perturbations causing the largest shifts in typical behavior often increased behavioral variability, which points to a trade-off we need to watch closely.
Rosa: So, the main implication is that we need a more comprehensive model card for VLA systems that reports on execution quality, not just success rates.
Dev: It suggests that if we want reliable physical deployment of these models, focusing on reducing variability in those key behavioral metrics under noise is a necessary next step.
Rosa: And so, today we've discussed the paper "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," exploring how to evaluate robustness beyond simple success rates and what that means for real-world deployment.
Dev: We'll be back next time with a new set of papers, so stay tuned.
Taro: I’m really looking forward to seeing how these behavioral metrics are applied in more complex autonomy scenarios.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications