Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration

arXiv:2608.06381 · cs.HC, cs.AI, cs.MA · Submitted 2026-06-09 · Read on arXiv

Mateus Levi Simões Fernandes, Alberto Sardinha

PUC-Rio

cs.HC, cs.AI, cs.MA

Submitted: 2026-06-09

Comments: 13 pages, 6 figures, accepted as an Extended Abstract at AAMAS 2026

Journal ref: Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems. 2026. p. 3259-3261

DOI: 10.65109/HDMG2174

Code: https://github.com/mateuslevisf/xai-hat-aamas2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 53/100

The gist: This paper addresses a "critical methodological gap" in human-agent teaming research: the fact that "no work has systematically evaluated XAI support generated from learned, high-performing policies

Terminology

Summary

This paper addresses a critical methodological gap in human-agent teaming research: the fact that no work has systematically evaluated XAI support generated from learned, high-performing policies in benchmark collaborative environments. Most existing research prioritizes optimizing agent performance over leveraging human understanding or relies on hand-crafted policies in a custom Minecraft environment, which limits transferability of their conclusions to state-of-the-art human-agent teaming research.

To address this, the researchers leverage the Hierarchical Ad Hoc Agents (HA2) architecture... which achieves state-of-the-art performance when collaborating with humans in Overcooked-AI. The HA2 architecture displays intrinsic interpretability because any action taken by the agent on the environment depends on comprehensible high-level subtask choices. The authors generate real-time explanations directly from the agent’s hierarchical decision-making process and deliver them through text or audio modalities via a novel trigger-based system.

Methodology

The study utilized a between-subjects experiment (n=38) with three conditions: Control (no explanations), Text, and Audio. Participants engaged in four 80-second gameplay sessions in Overcooked-AI’s Counter-Circuit layout. The explanation system uses a priority-based trigger system to ensure explanations are delivered only at coordination-relevant moments, such as Blocking (when a human occupies the agent's target tile), Critical Path (when the agent selects essential subtasks), or Subtask Change (when the Manager switches subtasks).

Results

  • Performance and Adaptation: The study found that simple status explanations provide no immediate performance benefit. In fact, the control condition scored higher (M=125.00, SD=18.19, n=14) than explanation conditions (M=114.17, SD=22.59, n=24), though this did not reach statistical significance. However, the researchers observed trends toward faster improvement rates, noting that Audio participants had the fastest score improvement (M=11.20 points/session, SD=6.94), followed by Text (M=8.00, SD=7.85), then Control (M=4.57, SD=8.96). This suggests XAI may enhance adaptation rather than immediate task execution.

  • Cognitive Workload: There were no significant cognitive load differences across conditions, suggesting that audio explanations... [may] reduce processing effort by making agent intentions immediately available without requiring visual attention shifts as with text.

  • Subjective Experience and Modality: A significant finding was that audio explanations produced a significant reduction in participants’ working-alliance bond with the agent – an effect absent under the text modality. Specifically, Audio participants rated the bond markedly lower (M = 3.23, SD = 0.50) than text (M = 4.26, SD = 1.11) or control (M = 4.65, SD = 0.97) on the 1–7 scale.

Discussion and Conclusion

The authors interpret the reduction in the working-alliance bond as the affective consequence of an unmissable claim-versus-reality gap inherent to model-free policy execution. They argue that spoken explanations activate partnership expectations the underlying reactive policy cannot meet, because the HA2 Manager... is model-free and selects subtasks reactively at each timestep; it cannot maintain the consistent, committal coordination that this framing implies. While text explanations can be ignored, audio cannot be ignored: it is continuously perceived and frames the agent as actively communicating intent.

The paper concludes that while hierarchical reinforcement learning architectures can serve as both high-performing agents and sources of authentic explanations without sacrificing either capability, effective collaborative XAI requires matching explanation richness to the underlying policy’s actual capacity for plan commitment and partner modelling. Future improvements should include extending the Manager with a commitment horizon longer than a single timestep and incorporating an explicit model of the human teammate.

Improvements for AI systems

1. Temporal Commitment-Aligned Policy (TCAP)

  • The Improvement: Integrate a temporal planning horizon into the Hierarchical Manager that synchronizes the agent's intent with its execution certainty.

  • What the improved system can do: The agent will only issue high-commitment, unmissable explanations (such as audio) when its internal policy has a mathematically verified commitment to a subtask for a specific duration. If the policy is highly reactive or uncertain, the system will automatically downgrade the explanation to a low-commitment modality (such as text or visual cues) to prevent the claim-versus-reality gap that damages human trust.

2. Uncertainty-Aware Modality Switching (UAMS)

  • The Improvement: Implement a decision engine that selects the communication modality (audio vs. text) based on the entropy (uncertainty) of the agent's current subtask selection.

  • What the improved system can do: The system will use audio for high-confidence, stable trajectories to minimize the human's cognitive load through hands-free perception. Conversely, when the agent's policy detects a high probability of a subtask change (due to environmental shifts or human movement), it will switch to text or visual overlays, signaling to the human that the information is provisional rather than committal.

3. Theory-of-Mind (ToM) Coordination Explanations

  • The Improvement: Augment the HA2 architecture with an explicit model of the human teammate’s predicted intent and state.

  • What the improved system can do: Instead of providing self-centric explanations (e.g., I am picking up an onion), the agent will provide relational explanations (e.g., I am picking up the onion because I see you are heading to the pot). This allows the agent to explain the reasoning behind the coordination rather than just the reasoning behind the action, directly addressing the need for partner modeling.

4. Adaptive XAI Scaffolding

  • The Improvement: Develop a closed-loop feedback mechanism that monitors the human partner's adaptation rate and performance trends.

  • What the improved system can do: The agent will provide high-frequency, multi-modal explanations during the initial learning phase of a human-agent pairing to accelerate the human's mental model formation. As the system detects that the human's performance has stabilized (indicating successful adaptation), it will automatically transition to a sparse explanation mode, delivering only critical, trigger-based information to prevent cognitive fatigue and over-reliance.

Abstract

Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research. We provide the first systematic evaluation of XAI support generated from an intrinsically explainable learned policy in an established benchmark. Using the Hierarchical Ad Hoc Agents (HA squared) architecture in Overcooked-AI, we generate real-time explanations from hierarchical subtask selections, delivered through text or audio via a novel trigger-based system. Our between-subjects experiment (n=38) found no significant performance effects, though participants with explanations showed trends toward faster performance improvement. More notably, audio explanations produced a significant reduction in participants' working-alliance bond with the agent -- an effect absent under the text modality -- suggesting that spoken explanations activate partnership expectations the underlying reactive policy cannot meet. We provide the first modality comparison in real-time human-agent collaboration and establish a baseline methodology for evaluating intrinsically explainable reinforcement learning architectures in benchmark environments. Results point to matching explanation modality to the underlying policy's capacity of sustaining the partnership its delivery implies as a potential path for more effective collaborative XAI.

Sources

Related papers