Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences
cs.AI, cs.HC
Submitted: 2025-02-02
Updated: 2026-09-08
Comments: EMNLP 2026
Code: https://github.com/google-research/scenic
License: http://creativecommons.org/licenses/by/4.0/
The gist: Natural-language instructions rarely specify every detail required for embodied action.
Terminology
Abstract
Natural-language instructions rarely specify every detail required for embodied action. An agent asked to ``prepare an apple,'' for example, must still determine whether to wash or cut it, where to place it, and in what order to perform these actions. Such decisions often reflect user-specific preferences that are demonstrated through behavior but never explicitly stated. We study whether embodied agents can infer these latent preferences from a small number of prior demonstrations and apply them when planning in new situations. To support systematic evaluation, we introduce Preference-based Planning (PbP), a benchmark containing 5,000 evaluation groups and 290 preferences organized into three levels: atomic action parameters, strategic interaction and placement policies, and temporal ordering constraints. We further propose Inferring the Unspoken (InTU), a two-stage framework that first verbalizes the preference inferred from multimodal behavioral demonstrations and then generates an action plan conditioned on that explicit representation. Experiments with video-language and language models reveal a substantial preference-acquisition gap: models plan effectively when given the ground-truth preference, but their performance degrades sharply when the same preference must be inferred from behavior. Explicit verbalization consistently improves alignment over direct end-to-end planning, particularly for strong multimodal models, and provides greater robustness when preferences must transfer across visually distinct scenes. These results identify visual-to-semantic preference acquisition, rather than preference-conditioned planning alone, as a central bottleneck in personalized embodied intelligence. They also demonstrate that language can serve as an interpretable and transferable intermediate representation between observed behavior and personalized action.
Sources
- GPT-4 Technical Report
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- On the Opportunities and Risks of Foundation Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
- AI2-THOR: An Interactive 3D Environment for Visual AI
- SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- LLaMA: Open and Efficient Foundation Language Models
- Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
- OPT: Open Pre-trained Transformer Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection