Infer Human's Intentions Before Following Natural Language Instructions
cs.AI, cs.CL, cs.LG
Submitted: 2024-09-26
Updated: 2026-08-26
Code: https://github.com/Simon-Wan/FISER
Project page: https://sites.google.com/view/fiser-hmt
License: http://creativecommons.org/licenses/by/4.0/
The gist: For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments.
Terminology
Abstract
For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the human speakers assume sufficient prior knowledge about their hidden goals and intentions. Standard language grounding and planning methods fail to address such ambiguities because they do not model human internal goals as additional partially observable factors in the environment. We propose a new framework, Follow Instructions with Social and Embodied Reasoning (FISER), aiming for better natural language instruction following in collaborative embodied tasks. Our framework makes explicit inferences about human goals and intentions as intermediate reasoning steps. We implement a set of Transformer-based models and evaluate them over a challenging benchmark, HandMeThat. We empirically demonstrate that using social reasoning to explicitly infer human intentions before making action plans surpasses purely end-to-end approaches. We also compare our implementation with strong baselines, including Chain of Thought prompting on the largest available pre-trained language models, and find that FISER provides better performance on the embodied social reasoning tasks under investigation, reaching the state-of-the-art on HandMeThat.
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- GOMA: Proactive Embodied Cooperative Communication via Goal-Oriented Mental Alignment
- COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection