Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
Matthew Francis Dixon
q-fin.RM, cs.AI, cs.LG, stat.ML
Submitted: 2026-06-16
Comments: 28 pages, 3 figures, 6 tables. Source code available from https://github.com/mfrdixon/agentic-AI-as-POMDP
Code: https://github.com/mfrdixon/agentic-AI-as-POMDP
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic artificial intelligence systems introduce a new class of model risk.
Terminology
Abstract
Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate forecasts, select actions, and adapt their behavior over time. Existing validation methodologies focus primarily on predictive accuracy and therefore provide limited insight into the quality of the underlying decision process. This paper proposes a model validation framework for agentic AI based on Partially Observable Markov Decision Processes (POMDPs). The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility, allowing each component to be validated independently. Large language models (LLMs) are formalized as approximate Bayesian filtering operators, and a model-risk taxonomy is developed encompassing state-space, filtering, forecast, policy, utility-specification, and parameter risks. The model risk validation methodology is demonstrated through a portfolio-management case study in which an agent infers latent market regimes from market and macroeconomic information, generates belief-conditioned forecasts, and constructs portfolios using a Black--Litterman framework. Empirical validation combines performance analysis, belief calibration diagnostics, coverage tests, ablation studies, and parameter-sensitivity analysis. The results indicate that latent-state inference contributes independently to decision quality and that the principal conclusions remain robust across a broad range of parameter values. The principal contribution of the paper is a practical framework for extending established model risk management concepts to autonomous AI systems and providing a rigorous foundation for their validation, governance, and monitoring.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- On the Opportunities and Risks of Foundation Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Holistic Evaluation of Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- The Rise and Potential of Large Language Model Based Agents: A Survey
Related papers
- Financial Tail Risk Beyond Lipschitz Continuity via Semi-Discrete Optimal Transport
- DisclosureBeta: A Measurement-Channel Theory for Regime-Conditioned Betas from LLM-Read Risk Disclosures
- Pricing the DeFi Tail: Do Protocols or Depositors Price Operational Risk?
- On the approximation of posterior laws in compound loss models by conditional Wasserstein GANs
- DTD-VAE: Disentangled Temporal Dependencies VAE for Credit Risk Prediction
- An Extreme Value Perspective on Learning Stress Laws