Coding Agents with Harness for Safe Robot Control
cs.RO, cs.AI, cs.CL, cs.CV
Submitted: 2026-09-17
Updated: 2026-10-02
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- WorldVLA: Towards Autoregressive Action World Model
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
- Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Agent as Policy for Robotic Manipulation
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation
- Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
- Code as Policies: Language Model Programs for Embodied Control
- Lost in the Middle: How Language Models Use Long Contexts
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving