Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation: A Case Study in Decision Support for Rice Cultivation in Japan
cs.AI, cs.HC
Submitted: 2025-12-24
Updated: 2026-09-19
Project page: https://christophm.github.io/interpretable-ml-book/Moshkovich
License: http://creativecommons.org/licenses/by/4.0/
The gist: Explainable artificial intelligence (XAI) reveals how explanatory variables relate to a response variable, yet communicating XAI outputs to laypersons remains difficult, limiting trust in AI-based
Terminology
Abstract
Explainable artificial intelligence (XAI) reveals how explanatory variables relate to a response variable, yet communicating XAI outputs to laypersons remains difficult, limiting trust in AI-based predictions. Large language models (LLMs) can translate technical explanations into accessible narratives, but iterative refinement of XAI explanations by an autonomous LLM agent remains unexplored. This study proposes an agentic XAI framework that combines SHapley Additive exPlanations (SHAP) with iterative refinement by a multimodal LLM and tests it as an agricultural recommendation system on rice yield data from 28 fields in Japan. From a SHAP result, the agent explored additional analyses across 11 refinement rounds (Rounds 0-10). Crop scientists (n = 12) and LLM judges (n = 14) scored every round on seven criteria: Specificity, Clarity, Conciseness, Practicality, Contextual Relevance, Cost Consideration, and Crop Science Credibility. Both groups found that refinement raised the average score by 30-33% over Round 0, peaking at Rounds 3-4, after which quality declined, below the starting point for crop scientists. Refinement therefore requires strategic early stopping, which challenges assumptions of monotonic improvement. Criterion-level trajectories indicate a bias-variance trade-off. Early rounds lacked Specificity (bias), whereas excessive iteration eroded Conciseness and raised Cost Consideration through ungrounded economic reasoning (variance). The LLM judges overscored every criterion by 1.4-2.3 points but largely preserved the experts' ranking of rounds (Spearman ρ = 0.58-0.90), so screened LLM judges can flag the quality peak despite unreliable absolute scores. Trustworthy agentic XAI also needs expert-anchored screening of LLM judges and transparent, verifiable refinement records.
Sources
- The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
- LLMs for Explainable AI: A Comprehensive Survey
- Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
- Towards A Rigorous Science of Interpretable Machine Learning
- Inherent and emergent liability issues in LLM-based agentic systems: a principal-agent perspective
- AI for a Planet Under Pressure
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Training Language Models to Self-Correct via Reinforcement Learning
- Agentic AI Process Observability: Discovering Behavioral Variability
- Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems
- Who is Responsible? The Data, Models, Users or Regulations? A Comprehensive Survey on Responsible Generative AI for a Sustainable Future
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- Farmer.Chat: Scaling AI-Powered Agricultural Services for Smallholder Farmers
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection