ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts
cs.AI, cs.RO
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/EmbodiedAISurvey/ConflictVLA-Bench
Terminology
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents
- HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
- SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
- Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing
- Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection