UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking
cs.RO, cs.CV, eess.IV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://tw5775.github.io/UniTrackPLA
Terminology
Sources
- Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking
- P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
- OmniVLN: Omnidirectional 3D Perception and Token-Efficient LLM Reasoning for Visual-Language Navigation across Air and Ground Platforms
- EAGOR: Embodied Reasoning in Omni-direction
- Hierarchical Instruction-aware Embodied Visual Tracking
- Follow-Bench: A Unified Motion Planning Benchmark for Socially-Aware Robot Person Following
- UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
- DINOv3
- Qwen3 Technical Report
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving