LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
cs.LG
Submitted: 2026-03-30
Updated: 2026-09-26
Comments: Accepted to EMNLP 2026 Main Conference (Oral Presentation). Code and benchmark: https://github.com/cau-hai-lab/LIBERO-Para
Code: https://github.com/cau-hai-lab/LIBERO-Para
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones.
Terminology
Abstract
Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data, leading to overfitting to specific instruction formulations and leaving robustness to paraphrased instructions underexplored. To study this gap, we introduce LIBERO-Para, a controlled benchmark that independently varies action expressions and object references for fine-grained analysis of linguistic generalization. Across seven VLA configurations (0.6B-7.5B), we observe consistent performance degradation of 22-52 pp under paraphrasing. This degradation is primarily driven by object-level lexical variation: even simple synonym substitutions cause large drops, indicating reliance on surface-level matching rather than semantic grounding. Moreover, 80-96% of failures arise from planning-level trajectory divergence rather than execution errors, showing that paraphrasing disrupts task identification. Binary success rate treats all paraphrases equally, obscuring whether models perform consistently across difficulty levels or rely on easier cases. To address this, we propose PRIDE, a metric that quantifies paraphrase difficulty using semantic and syntactic factors. Our benchmark and corresponding code are available at: https://github.com/cau-hai-lab/LIBERO-Para
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- Qwen3-VL Technical Report
- LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- OpenVLA: An Open-Source Vision-Language-Action Model
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- Decoupled Weight Decay Regularization
- Qwen2.5 Technical Report
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- GPT-4 Technical Report
- LIBERO-X: Robustness Litmus for Vision-Language-Action Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks