Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration
Mateus Levi Simões Fernandes, Alberto Sardinha
PUC-Rio
cs.HC, cs.AI, cs.MA
Submitted: 2026-06-09
Comments: 13 pages, 6 figures, accepted as an Extended Abstract at AAMAS 2026
Journal ref: Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems. 2026. p. 3259-3261
DOI: 10.65109/HDMG2174
Code: https://github.com/mateuslevisf/xai-hat-aamas2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 53/100
The gist: This paper addresses a "critical methodological gap" in human-agent teaming research: the fact that "no work has systematically evaluated XAI support generated from learned, high-performing policies
Terminology
Summary
This paper addresses a critical methodological gap
in human-agent teaming research: the fact that no work has systematically evaluated XAI support generated from learned, high-performing policies in benchmark collaborative environments.
Most existing research prioritizes optimizing agent performance over leveraging human understanding
or relies on hand-crafted policies in a custom Minecraft environment,
which limits transferability of their conclusions to state-of-the-art human-agent teaming research.
To address this, the researchers leverage the Hierarchical Ad Hoc Agents (HA2) architecture... which achieves state-of-the-art performance when collaborating with humans in Overcooked-AI.
The HA2 architecture displays intrinsic interpretability
because any action taken by the agent on the environment depends on comprehensible high-level subtask choices.
The authors generate real-time explanations directly from the agent’s hierarchical decision-making process and deliver them through text or audio modalities via a novel trigger-based system.
Methodology
The study utilized a between-subjects experiment (n=38)
with three conditions: Control (no explanations), Text, and Audio.
Participants engaged in four 80-second gameplay sessions in Overcooked-AI’s Counter-Circuit layout.
The explanation system uses a priority-based trigger system
to ensure explanations are delivered only at coordination-relevant moments,
such as Blocking
(when a human occupies the agent's target tile), Critical Path
(when the agent selects essential subtasks), or Subtask Change
(when the Manager switches subtasks).
Results
-
Performance and Adaptation: The study found that
simple status explanations provide no immediate performance benefit.
In fact,the control condition scored higher (M=125.00, SD=18.19, n=14) than explanation conditions (M=114.17, SD=22.59, n=24),
though thisdid not reach statistical significance.
However, the researchers observedtrends toward faster improvement rates,
noting thatAudio participants had the fastest score improvement (M=11.20 points/session, SD=6.94), followed by Text (M=8.00, SD=7.85), then Control (M=4.57, SD=8.96).
This suggestsXAI may enhance adaptation rather than immediate task execution.
-
Cognitive Workload: There were
no significant cognitive load differences across conditions,
suggesting thataudio explanations... [may] reduce processing effort by making agent intentions immediately available without requiring visual attention shifts as with text.
-
Subjective Experience and Modality: A significant finding was that
audio explanations produced a significant reduction in participants’ working-alliance bond with the agent – an effect absent under the text modality.
Specifically,Audio participants rated the bond markedly lower (M = 3.23, SD = 0.50) than text (M = 4.26, SD = 1.11) or control (M = 4.65, SD = 0.97) on the 1–7 scale.
Discussion and Conclusion
The authors interpret the reduction in the working-alliance bond as the affective consequence of an unmissable claim-versus-reality gap inherent to model-free policy execution.
They argue that spoken explanations activate partnership expectations the underlying reactive policy cannot meet,
because the HA2 Manager... is model-free and selects subtasks reactively at each timestep; it cannot maintain the consistent, committal coordination that this framing implies.
While text explanations can be ignored,
audio cannot be ignored: it is continuously perceived and frames the agent as actively communicating intent.
The paper concludes that while hierarchical reinforcement learning architectures can serve as both high-performing agents and sources of authentic explanations without sacrificing either capability,
effective collaborative XAI requires matching explanation richness to the underlying policy’s actual capacity for plan commitment and partner modelling.
Future improvements should include extending the Manager with a commitment horizon longer than a single timestep
and incorporating an explicit model of the human teammate.
Improvements for AI systems
1. Temporal Commitment-Aligned Policy (TCAP)
-
The Improvement: Integrate a temporal planning horizon into the Hierarchical Manager that synchronizes the agent's
intent
with itsexecution certainty.
-
What the improved system can do: The agent will only issue high-commitment,
unmissable
explanations (such as audio) when its internal policy has a mathematically verified commitment to a subtask for a specific duration. If the policy is highly reactive or uncertain, the system will automatically downgrade the explanation to a low-commitment modality (such as text or visual cues) to prevent theclaim-versus-reality
gap that damages human trust.
2. Uncertainty-Aware Modality Switching (UAMS)
-
The Improvement: Implement a decision engine that selects the communication modality (audio vs. text) based on the entropy (uncertainty) of the agent's current subtask selection.
-
What the improved system can do: The system will use audio for high-confidence, stable trajectories to minimize the human's cognitive load through hands-free perception. Conversely, when the agent's policy detects a high probability of a subtask change (due to environmental shifts or human movement), it will switch to text or visual overlays, signaling to the human that the information is
provisional
rather thancommittal.
3. Theory-of-Mind (ToM) Coordination Explanations
-
The Improvement: Augment the HA2 architecture with an explicit model of the human teammate’s predicted intent and state.
-
What the improved system can do: Instead of providing
self-centric
explanations (e.g.,I am picking up an onion
), the agent will providerelational
explanations (e.g.,I am picking up the onion because I see you are heading to the pot
). This allows the agent to explain the reasoning behind the coordination rather than just the reasoning behind the action, directly addressing the need for partner modeling.
4. Adaptive XAI Scaffolding
-
The Improvement: Develop a closed-loop feedback mechanism that monitors the human partner's adaptation rate and performance trends.
-
What the improved system can do: The agent will provide high-frequency, multi-modal explanations during the initial
learning phase
of a human-agent pairing to accelerate the human's mental model formation. As the system detects that the human's performance has stabilized (indicating successful adaptation), it will automatically transition to asparse explanation
mode, delivering only critical, trigger-based information to prevent cognitive fatigue and over-reliance.
Abstract
Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research. We provide the first systematic evaluation of XAI support generated from an intrinsically explainable learned policy in an established benchmark. Using the Hierarchical Ad Hoc Agents (HA squared) architecture in Overcooked-AI, we generate real-time explanations from hierarchical subtask selections, delivered through text or audio via a novel trigger-based system. Our between-subjects experiment (n=38) found no significant performance effects, though participants with explanations showed trends toward faster performance improvement. More notably, audio explanations produced a significant reduction in participants' working-alliance bond with the agent -- an effect absent under the text modality -- suggesting that spoken explanations activate partnership expectations the underlying reactive policy cannot meet. We provide the first modality comparison in real-time human-agent collaboration and establish a baseline methodology for evaluating intrinsically explainable reinforcement learning architectures in benchmark environments. Results point to matching explanation modality to the underlying policy's capacity of sustaining the partnership its delivery implies as a potential path for more effective collaborative XAI.
Sources
- A Hierarchical Approach to Population Training for Human-AI Collaboration
- On the Utility of Learning about Humans for Human-AI Coordination
- Collaborating with Humans without Human Data
- Explainable Artificial Intelligence: a Systematic Review
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support