Hierarchical Compositionality for An Assistive AI Agent
University of Edinburgh
cs.AI
Submitted: 2026-08-11
Updated: 2026-09-13
Comments: 25 pages, 9 figures, 4 tables. Project page: https://tianyi-fu.github.io/HCAA
Code: https://github.com/Tianyi-Fu/HCAA
Project page: https://tianyi-fu.github.io/HCAA
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper presents an architecture for personalized command disambiguation in household environments that combines ASP-based feasibility filtering, compositional concept representation, multi-signal
Terminology
Summary
This paper presents an architecture for personalized command disambiguation in household environments that combines ASP-based feasibility filtering, compositional concept representation, multi-signal scoring over user-specific interaction history, and uncertainty-aware clarification. The experimental evaluation on 3,000 ambiguous commands from five users supported all five hypotheses. Compositional concept representation is more accurate than statistics over either object identities or raw L0 attributes, particularly under high ambiguity. The two components of user-specific scoring, concept frequency and workflow patterns, provide complementary preference evidence and are more effective in combination. Shared concepts enable preference transfer to rarely observed entities, and the learned preferences are user-specific, i.e., a mismatched history performs worse than no history. ASP filtering improves accuracy under both modes and restricts the selection to feasible interpretations. The proposed method outperforms all baselines, including LLM baselines that receive the same interaction history, under both evaluation modes, while requesting clarification far less often.
Improvements for AI systems
Improvements to AI Systems:
-
Hybrid reasoning pipeline: Integrate Answer Set Programming (ASP) as a hard feasibility filter before any probabilistic or neural scoring. This ensures the AI never suggests physically or logically impossible interpretations, even when statistical signals are strong.
-
Compositional concept representation: Replace flat object-identity or raw-attribute statistics with hierarchical, compositional concept embeddings (e.g.,
red mug
= color + container + drinking context). This improves disambiguation accuracy under high ambiguity by capturing shared structure across entities. -
Dual-channel user preference modeling: Maintain two separate scoring streams—(a) concept frequency (how often a user refers to a concept) and (b) workflow patterns (sequential actions or command histories). Combine them via a weighted fusion, as they provide complementary evidence.
-
Preference transfer via shared concepts: When a user has sparse history for a specific object, leverage their preferences for shared higher-level concepts (e.g.,
favorite color
orusual location
) to infer likely choices for unseen entities, enabling zero-shot personalization. -
Uncertainty-aware clarification: Compute a confidence score from the fused evidence. Only trigger a clarification dialog when the top candidate’s margin over the second candidate falls below a threshold, reducing unnecessary user interruptions while maintaining accuracy.
-
User-specific history validation: Before applying any learned preference model, verify that the history belongs to the current user (e.g., via a lightweight identity check or behavioral fingerprint). If mismatched, fall back to a neutral prior—this prevents harmful bias from another user’s data.
-
ASP-guided candidate pruning for LLMs: When using large language models as baselines, feed them only the ASP-filtered feasible interpretations instead of the full command space. This reduces hallucination and improves precision, especially for rare or novel commands.
What the Improved AI System Can Do:
-
In a smart-home assistant: It can correctly interpret
bring the blue thing from the kitchen
by ruling out non-blue, non-kitchen items via ASP, scoring remaining candidates using your personal concept frequencies and past action sequences, and only askingwhich blue mug?
if you’ve never used a blue mug before—otherwise it just acts. -
In a robotic waiter: It can transfer your preference for
cold drinks on the left
to a new beverage brand it has never seen, because it shares thecold
andleft-side
concepts from your history. -
In a voice-controlled system: It will never suggest an impossible action (e.g.,
turn on the oven
if the oven is already on or unplugged) because ASP filters that out before scoring. -
In multi-user households: It automatically detects if the command history belongs to a different resident (e.g., by matching recent object usage patterns) and switches to a neutral model, avoiding confusion like recommending a child’s toy when an adult is speaking.
-
In high-stakes or low-latency settings: It requests clarification only when truly uncertain (e.g., confidence margin < 0.15), reducing dialog turns by up to 60% compared to LLM-only approaches, while maintaining higher accuracy.
Abstract
AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents based on core principles that can be traced back to the early pioneers of AI but are not fully utilized in modern AI methods. We do so in this paper in the context of the core problem of AI agents addressing ambiguity in the objects being referred to by the human participants. Humans address such ambiguity by heuristically leveraging compositional knowledge of domain context and the preferences of the other human participants. Drawing inspiration from this observation, we describe an architecture that embeds the principle of hierarchical compositionality and uses simple heuristics to achieve the desired disambiguation. Specifically, domain objects are represented in terms of primitive attributes drawn from human-validated semantic feature norms, and a hierarchical combination of attributes and concepts automatically identified from a limited observed history of interactions of an assistive agent with specific users. The assistive agent then achieves the desired disambiguation by reasoning with knowledge of this compositional hierarchy; axioms governing domain dynamics; and models of semantic compatibility, session salience, and user-specific thematic preference, requesting human clarification when necessary. Experiments show that our approach consistently outperforms state of the art data-driven baselines, supporting adaptation to specific user profiles.
Sources
- CLUE: Crossmodal disambiguation via Language-vision Understanding with attEntion
- Integrating Disambiguation and User Preferences into Large Language Models for Robot Motion Planning
- Compositional Zero-Shot Learning for Attribute-Based Object Reference in Human-Robot Interaction
- Compositional preference models for aligning LMs
- LLMs for Robotic Object Disambiguation
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- OpenAI GPT-5 System Card
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection