Autonomous Assessment of Generalizability of AI Agent Capabilities

arXiv:2512.16733 · cs.AI · Submitted 2025-12-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Autonomous Assessment of Generalizability of AI Agent Capabilities".

Jane: The paper was written by Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi et al. from Arizona State University and Indian Institute of Technology Madras.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, you have to see the title of this new paper we just got from the arXiv feed.

Jane: Is it "Autonomous Assessment of Generalizability of AI Agent Capabilities"?

Tom: That is exactly the one!

Jane: It sounds quite intimidating at first glance, doesn't it?

Tom: It really does, but it’s basically asking how we can automatically check if an AI agent can actually do what we think it can do in different situations.

Jane: Right, because right now, we often just hope the agent doesn't fail when things change slightly.

Lu: Imagine a robot sent to a house it has never seen before to find a lost cat.

Jane: That’s a perfect way to put it, Lu!

Lu: If we can't autonomously assess if that robot understands "find the cat" across different layouts, we're just playing a dangerous game of chance.

Meng: I struggle with that in my daily work because testing every single edge case manually is impossible for any engineer.

Tom: That’s why this paper from the teams at Arizona State University and IIT Madras is such a big deal, Meng.

Meng: They're trying to automate the very process of finding where the agent breaks.

Jane: And that's a huge shift from just training an agent to actually auditing it.

Lalam: This moves us toward a culture of accountability where we don't just trust an AI because it looks smart in a demo.

Tom: Exactly, Lalam, we need to know the limits before we let them loose in the real world.

Jane: It's about building a bridge of reliability between the code and the user.

Tom: So, how exactly do they go about this "autonomous assessment" they're promising?

Summary: Tom: We just talked about why we need this, but now let's look at how they actually built it.

Jane: They’ve introduced something called Monte Carlo Query Synthesis, or MCQS for short.

Tom: It sounds like a mouthful, but the concept is brilliant.

Jane: Think of it like a highly intelligent interrogator who doesn't just ask random questions, but specifically asks the ones most likely to make a suspect slip up.

Lu: I love that analogy!

Lu: Instead of just wandering around a simulator, the system uses Monte Carlo Tree Search to pick the most "informative" actions to try.

Meng: Does that mean it's actively looking for the gaps in its own knowledge?

Jane: Precisely, Meng.

Jane: It maintains two different versions of what it thinks the agent can do—one very pessimistic and one very optimistic.

Tom: And then it looks for a way to prove which version is right!

Meng: So it's basically playing a game of "prove me wrong" against itself?

Jane: That's a great way to describe it.

Jane: It uses the difference between those two extreme models to decide what the next test should be.

Lu: This creates a feedback loop where the system gets smarter about its own ignorance very quickly.

Lalam: It's like an AI learning to teach itself how to be safe, which is a profound leap for machine autonomy.

Tom: It really is, and it makes me wonder what kind of weird behaviors they actually caught using this method.

Improvements: Tom: We've seen the theory, but the actual results in this paper are where things get really interesting.

Jane: They compared MCQS to just testing things at random, and the difference is massive.

Tom: It’s not even a close race!

Jane: The paper shows that MCQS reaches a high level of accuracy with far fewer steps in the simulator than random testing ever could.

Meng: That's huge for us because simulator time and compute costs are incredibly expensive in production.

Lu: But the most shocking part was what they actually found out about these "smart" agents.

Jane: Oh, you mean the "superstitious" behaviors?

Lu: Yes! They found that an agent like GPT-five point four would sometimes open a completely unnecessary door just because it happened to be near its goal.

Tom: That is so bizarre! It's like the AI developed a weird habit that doesn't actually help it.

Jane: And they found even more with SayCan, where it struggled to stack blocks because it would misidentify colors.

Meng: It shows that even if the logic is there, the execution can be incredibly fragile.

Lu: It makes me wonder what else is hiding in the shadows of these foundation models.

Lalam: If we can use this to clean up these "superstitions," we'll create machines that feel much more natural and predictable to humans.

Tom: This paper really changes how we look at agent reliability, doesn't it?

Conclusion: Tom: Well, we have covered a lot of ground today on the paper "Autonomous Assessment of Generalizability of AI Agent Capabilities."

Jane: It’s clear that moving from manual testing to this kind of active, automated auditing is the next frontier.

Tom: Before we head out, I want to hear one last thought from the team.

Lu: I see a future where every autonomous system carries its own "stress-test" engine to constantly re-evaluate itself in new environments.

Meng: From my side, I just want to see how we can integrate this into standard CI/CD pipelines so safety is baked in from day one.

Lalam: And I believe this will ultimately foster a much deeper sense of social trust as AI becomes a seamless part of our daily lives.

Jane: Thanks for joining us, everyone!

Tom: We'll see you next time with another deep dive into the latest research!

Arizona State University · Indian Institute of Technology Madras

cs.AI

Submitted: 2025-12-18

Updated: 2026-09-14

Code: https://github.com/tomsilver/pddlgym

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 80/100

The gist: This paper introduces Monte Carlo Query Synthesis (MCQS), an active query-synthesis method designed to learn symbolic stochastic capability models of black-box AI (BBAI) systems.

Key concepts

Monte Carlo Query Synthesis (MCQS)
An auditing method that uses Monte Carlo Tree Search to select the most informative actions for testing an AI agent. It compares pessimistic and optimistic models of an agent's capabilities, using their differences to decide which tests will best reveal the agent's true limits and knowledge gaps.
Generalizability
The ability of an AI agent to perform its intended tasks successfully across different environments or situations. Instead of just working in a specific demo, a generalizable agent can adapt when layouts or conditions change slightly, ensuring it doesn't fail unexpectedly in new settings.
Superstitious Behaviors
Bizarre, unnecessary habits developed by AI agents that do not actually help achieve a goal. Examples include an agent opening an irrelevant door simply because it is near its target or misidentifying colors, which demonstrates that even logically sound models can have fragile execution.

Terminology

Summary

This paper introduces Monte Carlo Query Synthesis (MCQS), an active query-synthesis method designed to learn symbolic stochastic capability models of black-box AI (BBAI) systems. As foundation-model agents are increasingly deployed for sequential decision-making, characterizing their capabilities—what they can do, when they can do it, and what outcomes may result—is essential for safe deployment and avoiding unintended superstitious behaviors.

The Capability Modeling Framework

The authors formalize capability discovery as an active learning problem over a hypothesis class of capability models. Each model induces a distribution over state-trajectories resulting from the agent's execution of policies. These models are represented using conditional probabilistic effects, where:

  • Conditions may include conjunctions and disjunctions of literals in an interpretable symbolic space.

  • Outcomes are stochastic, modeled as probability distributions over effects.

  • The model captures that outcome distributions depend on the initial state and can involve non-liftable asymmetries.

By using this framework, the goal is to minimize the variational distance from the unknown true model h*, providing interpretable models that support precise guarantees of soundness and accuracy.

How MCQS Works

MCQS addresses the combinatorial challenge of searching through policy and hypothesis spaces by leveraging lattice theory. The set of hypotheses consistent with observed transitions forms a subset lattice, allowing the algorithm to track only two extremal elements:

  1. The pessimistic model (h), which is the lattice-meet representing the most conservative hypothesis consistent with observations.

  2. The optimistic model (h), which is the join of all models that are possibly sound w.r.t D.

The algorithm identifies discriminating policies by formulating a distinguishing Markov decision process (MDP). It uses Monte Carlo Tree Search (MCTS) to synthesize queries that maximize the expected divergence between trajectory distributions induced by these two competing hypotheses.

Query Synthesis Variants

The paper implements two distinct versions of the algorithm to handle different computational requirements:

  • MCQS-Exact (MCQS-E): This variant solves the distinguishing MDP directly by maintaining explicit predicted distributions over represented states using bit-vectors. It uses total variation distance as a reward to prioritize policies that assign mass to states likely under one model but unlikely under the other.

  • MCQS with set-of-support (SoS) factorization (MCQS-S): This sample-based approximation propagates only the intersection of the two supports of the predicted distributions. It utilizes a symmetric difference measure to focus on states that are deterministically falsifying, meaning observing such a state immediately rules out one of the hypotheses.

Empirical Findings and Limitations

Experiments across diverse agents—including GPT-5.4, CSteve, SayCan, and LAO*—demonstrate that MCQS learns accurate models more efficiently than random querying. The method successfully uncovered several salient capabilities and surprising limitations:

  • Superstitious behaviors, such as an agent opening an unnecessary door when tasked with retrieving a key.

  • Success probabilities that drop significantly under specific conditions, such as when an agent is holding a key.

  • High stochasticity in robotic agents like SayCan, where desired effects occurred with vanishingly low probability.

The authors conclude that because emerging agent-capabilities often do not exhibit symmetries over object substitution, learned models must be grounded rather than lifted to maintain modeling fidelity.

Improvements for AI systems

1. Automated Capability Auditing Module (Pre-Deployment Stress Testing)

  • What it can do: This module replaces generic unit testing with active, adversarial discovery of agent limitations. By utilizing Monte Carlo Query Search (MCQS), the system automatically synthesizes high-information queries designed to find edge cases—specifically targeting states where the agent’s behavior is most uncertain or likely to produce unintended side effects (e.g., superstitious behaviors like opening unnecessary doors). It provides a rigorous, symbolic report of an agent's success probabilities and side-effect distributions across all critical intents before the agent is ever exposed to a production environment.

2. Symbolic Probabilistic Capability Dashboard (User-Facing Interpretability)

  • What it can do: This transforms the black-box experience into a transparent, risk-aware interface. Instead of receiving a binary success/fail response, users are provided with an interpretable symbolic model of the agent's capabilities. For any given command, the system will output explicit conditional probabilities (e.g., Intent: Retrieve Key Success Probability: 0.85 Side Effect [Unnecessary Door Open]: 0.42 Required Condition: Agent must be in SW quadrant). This allows human operators to make informed, high-stakes decisions based on quantified risk rather than intuition.

3. Targeted Capability-Gap Fine-Tuning (Closed-Loop Training Optimization)

  • What it can do: This uses the information-rich state-action trajectories discovered by MCQS to optimize the training pipeline. Rather than wasting compute on redundant data, this system identifies the exact state-space regions where the agent's learned capability model deviates from its intended behavior. These high-divergence trajectories are then fed into Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) loops to specifically train the agent to resolve identified failures, such as misidentifying objects or failing in specific spatial configurations.

4. Runtime Risk-Aware Policy Monitor (Real-Time Safety Guardrails)

  • What it can do: This integrates the learned symbolic capability models into a real-time execution monitor. As an agent generates a multi-step plan, the monitor simulates the predicted outcome distribution using the MCQS model. If the predicted trajectory exceeds a predefined risk threshold—such as a high probability of an unintended side effect or a drop in success probability due to environmental stochasticity—the monitor intercepts the command and triggers an automatic re-planning request or demands human intervention before the agent executes the risky action.

Abstract

Safe deployment of black-box AI (BBAI) systems such as foundation model agents requires methods for evaluating their capabilities in novel settings. We define an agent's capability as its ability to achieve a short term objective and formalize the problem of learning models that predict whether, with what effects, and under what conditions, an agent can perform a capability. We introduce Monte Carlo Query Search (MCQS), an active query-synthesis method for learning symbolic stochastic capability models of BBAIs. MCQS models capabilities as conditional probability distributions over outcomes and formulates capability evaluation as an active learning problem over policies. We use Monte Carlo tree search to synthesize queries that maximally distinguish between extremal capability hypotheses: the lattice meet and join corresponding to the most pessimistic and optimistic models consistent with observed behavior. Executing these queries yields trajectories that prune inconsistent hypotheses. We prove soundness, completeness, and convergence properties under standard realizability and sampling assumptions. Experiments with multiple BBAI systems show that MCQS learns accurate capability models more efficiently than baseline query strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.

Sources

Related papers