Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks".
Dev: Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're discussing the title and authors of "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," and what that really implies for our work.
Dev: The authors are Higham, Izzo, Matteucci, and Suglia, and the paper essentially asks if just achieving a task success rate is sufficient when we consider real-world disturbances.
Taro: I think it's important because they’re pushing the idea that performance in the lab doesn't automatically translate to reliable behavior when environmental conditions shift unexpectedly.
Rosa: That’s right, Taro, and they are proposing this benchmark-agnostic evaluation framework to characterize how successful trajectories behave under perturbation.
Dev: It seems like their main point is that we need metrics that capture the nature of the execution—like how smooth or fast it was—instead of just a single success score.
The paper's summary: Rosa: So, to summarize what they’ve done in "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," they took existing benchmarks like LIBERO and LIBERO-Plus and extended them to test three state-of-the-art VLA models.
Dev: They evaluate these models across four task suites and seven different perturbation conditions, looking at both the typical successful behavior and how variable that behavior is.
Taro: They are calculating metrics like duration, total gripper movement, Cartesian jerk, and joint jerk to get a richer picture of the robot's execution style.
Rosa: They found that perturbations can definitely change the behavior of those successful trajectories, which is something TSR doesn't capture on its own.
Dev: Specifically, they calculated "task-level percentage change in the mean metric value" relative to a no-perturbation baseline for all these metrics.
The paper's improvements: Rosa: Now, regarding the improvements they suggest for this work, it seems they are moving away from relying only on TSR and instead proposing a suite of behavioral metrics that characterize motion smoothness, efficiency, and gripper behavior.
Dev: They propose calculating things like the duration in seconds to isolate task completion time from inference latency or measuring total gripper movement in meters to capture all the open and close motion.
Taro: I see them focusing on jerkiness, both Cartesian and joint jerk, which should give us a good measure of trajectory smoothness under stress.
Rosa: And they also look at path lengths—Cartesian path length for how far the end-effectors move in physical space, and joint path length for how much the arm configuration changes over time.
Dev: They quantify this by calculating both mean and P95 versions of those jerk metrics to see not just the average behavior but also the upper tail of what happens during execution.
Conclusion: Rosa: So, to wrap things up on "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," they conclude that successful trajectories are not behaviorally invariant to perturbations, showing that the model-dependent sensitivity is quite pronounced.
Dev: They also found correlations between TSR degradation and various behavioral metrics, like total gripper movement and mean Cartesian jerk, suggesting success isn't fully capturing execution quality.
Taro: And interestingly, they noted that the perturbations causing the largest shifts in typical behavior often increased behavioral variability, which points to a trade-off we need to watch closely.
Rosa: So, the main implication is that we need a more comprehensive model card for VLA systems that reports on execution quality, not just success rates.
Dev: It suggests that if we want reliable physical deployment of these models, focusing on reducing variability in those key behavioral metrics under noise is a necessary next step.
Rosa: And so, today we've discussed the paper "Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks," exploring how to evaluate robustness beyond simple success rates and what that means for real-world deployment.
Dev: We'll be back next time with a new set of papers, so stay tuned.
Taro: I’m really looking forward to seeing how these behavioral metrics are applied in more complex autonomy scenarios.
Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia
University of Edinburgh · Politecnico di Milano
cs.RO
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/esgi-research-group/vla-reliability
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 86/100
The gist: Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks, but this work proposes a benchmark-agnostic evaluation framework to measure behavioral
Key concepts
- Task Success Rate (TSR)
- A measure of whether a robot successfully completes its assigned manipulation task. It's the standard way to judge if an AI model achieved its goal, but this paper argues it is insufficient because it doesn't reveal *how* the success was achieved.
- Behavioral Robustness
- The ability of a VLA model to maintain consistent and similar execution metrics (like movement smoothness or trajectory path) even when the input data is slightly altered by noise or perturbations. This measures how stable the robot's actual movements are under stress.
- Trajectory-level Metrics
- Specific quantitative measurements taken from a successful robot movement, such as Cartesian jerk (how fast the end-effector accelerates), total gripper movement, and path length. These metrics describe the physical quality and style of the successful execution.
- Behavioral Variability (MAD)
- A statistical measure used to quantify how much a set of successful movements differs from one another under a specific condition. High variability suggests that while success is achieved, the robot's execution style is inconsistent.
Terminology
Summary
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks, but this work proposes a benchmark-agnostic evaluation framework to measure behavioral robustness by characterizing how successful trajectories are executed under input perturbations. The core finding is that perturbations can alter the behavior of successful trajectories, a phenomenon not captured by Task Success Rate (TSR) alone.
How it works
The proposed methodology extends widely-used benchmarks like LIBERO and LIBERO-Plus to evaluate the behavioral robustness of state-of-the-art VLA models. The framework moves beyond binary task success by calculating metrics that characterize the nature and variability of successful task execution, complementing TSR. The evaluation involves analyzing model behavior under original, unperturbed tasks as a baseline and under various perturbation conditions.
The specific behavioural evaluation metrics include:
-
Duration (s): Calculated as elapsed simulator time between the first and final recorded control step using the simulator control frequency to isolate task completion time from inference latency.
-
Total gripper movement (m): The cumulative absolute change in gripper width for a symmetric two-finger gripper, capturing total open/close motion.
-
Cartesian jerk (m/s3): Computed as the third derivative of the Cartesian position, capturing how abruptly the end-effector motion changes and thus trajectory smoothness. Mean and P95 versions are calculated to capture average and upper-tail behavior, respectively.
-
Joint jerk (rad/s3): Captures the smoothness of movement of the arm as a whole.
-
Cartesian path length (m) and Joint path length (rad): These metrics capture how far the end-effectors move through physical space and how much the robot arm configuration changes over time, respectively.
Quantifying typical behaviour change
To quantify how successful trajectories change under perturbation, the researchers calculate a distribution of trajectory-level metric values for both baseline and perturbed conditions. For each task, metric, and perturbation condition, they compute the mean metric value across the distribution of successful trajectories (Equation 5). This yields a task-level percentage change in the mean metric value
relative to the no-perturbation baseline. The overall effect of a perturbation across an entire task suite is summarized by reporting the median of these task-level percentage changes, ∆e m,p = mediank (∆k,m,p)
(Equation 6). Statistical significance is assessed using a two-sided Mann-Whitney U test with Benjamini-Hochberg false discovery rate correction.
Quantifying behaviour variability
In addition to characterizing typical behavior change, the framework quantifies behavioral variability using the median absolute deviation (MAD). For each task, metric, and condition, MAD is calculated as the median absolute deviation of trajectory-level metric values (Equation 7). The percentage change in MAD relative to the baseline is then calculated as ∆MADk,m,p (Equation 8). Positive values indicate increased behavioral variability. The median of these task-level percentage changes across all tasks in a suite is reported to summarize variability effects.
Key findings on behaviour and TSR correlation
The study investigates three research questions: RQ1 (How does typical behavior change?), RQ2 (Are perturbation-induced behavioral changes reflected by changes in the TSR?), and RQ3 (Do perturbations alter the behavioral consistency of successful trajectories?).
Regarding typical behavior, the results indicate that successful trajectories are not behaviourally invariant to perturbations.
For instance, notable examples from the Spatial suite showed a 24.1% (7/10 tasks significant) and 35.3% (10/10 tasks significant) median increase in mean Cartesian jerk under the camera perturbation
for the π0.5 model. This sensitivity is model-dependent; The robot and camera perturbations appear to have the most significant effect on the behaviour metrics of π0.5 and VLANeXt models.
Concerning TSR correlation, a Spearman rank correlation analysis revealed moderate positive correlations between TSR degradation and total gripper movement (ρ=0.56), P95 Cartesian jerk (ρ=0.54), duration (ρ=0.52), mean Cartesian jerk (ρ=0.43) and mean joint jerk (ρ = 0.4).
This suggests that perturbation conditions associated with larger reductions in task success also tended to produce larger behavioral changes in successful trajectories, indicating that changes in task success do not fully capture changes in successful execution behaviour.
Key findings on behavioural consistency
The analysis of behavioral consistency shows an interesting pattern: the perturbation conditions which produced the largest shifts in typical successful behaviour also frequently increased behavioural variability (Table III).
Notably, the camera perturbation led to an increase in MAD for every evaluated model and suite across duration, mean Cartesian jerk and gripper movement. Conversely, several perturbations led to reduced MAD when compared with the unperturbed baseline,
such as background, language, lighting and object conditions. This suggests that perturbations might be narrowing the set of behaviours that enable success.
Improvements for AI systems
Based on the scientific paper, here are specific improvements that can be made to existing Vision-Language-Action (VLA) models and what those improved systems could achieve:
-
Improve model evaluation by moving beyond Task Success Rate (TSR) as the sole metric for robustness.
-
Implement a benchmark-agnostic framework that complements TSR with trajectory-level behavioral metrics, specifically focusing on motion smoothness, efficiency (duration), and gripper behavior (total movement).
-
Analyze how perturbations change the mean and variability of these specific behavioral metrics during successful task execution (RQ1).
-
Quantify the correlation between TSR degradation and changes in these behavioral metrics to determine if success is decoupled from good execution quality (RQ2).
-
Assess whether different models exhibit differential robustness—i.e., some models may maintain high TSR while their successful trajectories suffer substantial behavioral degradation under the same perturbation (RQ2).
-
Determine if perturbations alter the consistency of successful behaviors by calculating and analyzing the Median Absolute Deviation (MAD) for these metrics, identifying which perturbations lead to more or less reliable execution strategies (RQ3).
This improved AI system would be able to:
-
Identify VLA models that achieve high task success rates but execute those successes in a
brittle
or inconsistent manner when faced with real-world environmental variations (e.g., changes in lighting, camera angles, or robot state). -
Provide a more comprehensive
model card
for VLA systems by reporting not just whether they succeed, but also the quality of their successful actions (e.g., smoothness and energy efficiency). -
Diagnose precisely which types of input perturbations (camera changes vs. language instruction vs. lighting) cause specific degradation in robotic movement patterns during task completion.
-
Help developers understand the trade-off between achieving a high TSR and maintaining a predictable, smooth, and efficient execution profile for physical deployment in real-world scenarios.
-
Guide the fine-tuning process by suggesting methods to improve trajectory consistency (reduce MAD) under challenging conditions to ensure reliable operation in variable environments.
Sources
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Colosseum V2: Benchmarking Generalization for Vision-Language-Action Models
- RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- OpenVLA: An Open-Source Vision-Language-Action Model
- MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
- vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
- Do World Action Models Generalize Better than VLAs? A Robustness Study
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving