Cognitive Action Reasoning for Proactive Robots from Human-Centered Multimodal Observations
cs.RO, cs.CV
Submitted: 2026-09-20
Updated: 2026-09-20
Terminology
Sources
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Large Vision-Language Models as Emotion Recognizers in Context Awareness
- LiMoDE: Rethinking Lifelong Robot Manipulation from a Mixture-of-Dynamic-Experts Perspective
- AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
- EmpathyAgent: Can Embodied Agents Conduct Empathetic Actions?
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- ProAct: A Dual-System Framework for Proactive Embodied Social Agents
- Event-Driven Proactive Assistive Manipulation with Grounded Vision-Language Planning
- Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data
- SASI: Leveraging Sub-Action Semantics for Robust Early Action Recognition in Human-Robot Interaction
- GPT-4 Technical Report
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Qwen2.5-Omni Technical Report
- GPT-4o System Card
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving