Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle
cs.AI, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen.
Terminology
Abstract
Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen. However, LLMs often generate plausible but incorrect statements, and no existing approach systematically tests whether such explanations are faithful to the underlying environment. Two classic software testing challenges stand in the way: there is no oracle for the correctness of an explanation, and the test inputs, natural language queries about a policy's behavior, lack the structure needed for systematic test case generation. We address both. Probabilistic model checking provides the test oracle, computing exact reference results against which LLM answers are graded automatically. A taxonomy of post hoc query categories structures the input space around the environment-level facts from which policy explanations are composed; test cases generated from it are prioritized by question-specific diagnostic difficulty scores. Across seven MDP environments, the testing separates three open-weight LLMs: a reasoning model passes 85% of test cases, a mid-size model 70%, and a 1B model falls below the random baseline, while prioritization surfaces significantly harder cases than random selection. Our results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.
Sources
- LLMs for Explainable AI: A Comprehensive Survey
- Formally Verifying and Explaining Sepsis Treatment Policies with COOL-MC
- Probabilistic Model Checking of Stochastic Reinforcement Learning Policies
- Verifying Memoryless Sequential Decision-making of Large Language Models
- TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models
- In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
- On the Fundamental Limits of LLMs at Scale
- Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments
- A Survey of Explainable Reinforcement Learning: Targets, Methods and Needs
- Gemma 3 Technical Report
- Gemma 4 Technical Report
- Model-Agnostic Policy Explanations with Large Language Models
- Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
- Explaining Agent Behavior with Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection