Instance-anchored interaction evidence: Grounding robot plans in human pointing and handling
cs.RO
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/WeiZhou96/iae-watchact
Terminology
Sources
- WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation
- GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
- VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
- Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
- Can Vision-Language Models Solve the Shell Game?
- GIVE: Grounding Human Gestures in Vision-Language-Action Models
- GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
- Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
- VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
- Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
- XSkill: Cross Embodiment Skill Discovery
- Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
- EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
- DeePoint: Visual Pointing Recognition and Direction Estimation
- YouRefIt: Embodied Reference Understanding with Language and Gesture
- Understanding Embodied Reference with Touch-Line Transformer
- Gesture-Informed Robot Assistance via Foundation Models
- Communicating human intent to a robotic companion by multi-type gesture sentences
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving