Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation
cs.RO
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/robocurve/inspect-robots
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Qwen3-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- SAM 3: Segment Anything with Concepts
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Show-Harness: Just a VLM Agent Can Play Robots
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- LoRA: Low-Rank Adaptation of Large Language Models
- VIA: Visual Interface Agent for Robot Control
- RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- TAPNext++: What's Next for Tracking Any Point (TAP)?
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- Code as Policies: Language Model Programs for Embodied Control
- Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness
- GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving