Using Grounded Theory for Agent Behavior Analysis at Scale

arXiv:2608.30391 · cs.CL, cs.AI · Submitted 2026-08-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Using Grounded Theory for Agent Behavior Analysis at Scale".

Jane: ===CATEGORIES===

"category name": "Difficulty in task management", "definition": "This captures instances where the sequence of steps or required inputs leads to procedural failure.", "member codes": ["trajectory id": "T101", "code": [Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about who put this paper together. The authors are Zhuoran Lu, Yangyang Yu, Zhuoyan Li, Yibo Meng, Nan Jiang, Chengxi Zang, Jie Gao, and Ziang Xiao from Purdue University and Stevens Institute of Technology.

Jane: That’s a solid group of researchers bringing different backgrounds to the table for this kind of interdisciplinary work. It shows how many different fields are needed to tackle these complex agent problems effectively.

Lu: Lu is clearly driving this, given the focus on applying qualitative methods to agent behavior analysis at scale, which requires a strong foundation in both computer science and social sciences.

Meng: Meng’s involvement suggests they’re thinking about how this theory actually translates into something we can build or implement in a practical setting.

Lalam: It's impressive that such diverse expertise is coming together to tackle something as abstract as agent behavior analysis. That kind of collaboration fuels real innovation.

The paper's summary: Tom: Now, let’s look at what the paper actually proposes in "Using Grounded Theory for Agent Behavior Analysis at Scale." Essentially, they introduce AutoTraceGT, which is a multi-agent pipeline that automates grounded theory on agent trajectories.

Jane: So it’s not just a description of behavior; it's an automated method for discovering those behaviors by iteratively performing open, axial, and theoretical coding until saturation is reached.

Lu: That iterative process, using constant comparison against a running codebook driven by strategic sampling, is where the core methodology lives. It’s essentially letting the data build the taxonomy itself.

Meng: If it automatically generates a behavioral taxonomy tailored to each specific task, that could drastically reduce the manual effort needed to categorize agent failures or successes in testing environments.

Lalam: Imagine how this could improve our internal understanding of AI interaction; instead of us guessing what goes wrong, we get a structured narrative derived directly from the agent's actions.

The paper's improvements: Tom: The authors highlight some specific improvements in AutoTraceGT, particularly regarding its reliability and quality. They show that across multiple datasets and different backbone LLMs, the method drives codebooks toward saturation reliably.

Jane: They also found that the induced codebooks cover a large percentage of failure modes found in human-annotated taxonomies while also surfacing patterns those traditional taxonomies miss entirely.

Lu: That finding about surfacing patterns missed by existing human analyses is significant because it suggests the method can uncover novel ways agents fail that we hadn't considered before.

Meng: From an engineering perspective, if the codebook can be repurposed as a deductive feature space for failure prediction, that opens up new avenues for building more robust diagnostic tools.

Lalam: That ability to create a feature space based on emergent theory is fascinating; it moves us beyond simply classifying known errors toward understanding the underlying principles of agent decision-making itself.

Conclusion: Tom: To wrap things up, the core implication of this paper, "Using Grounded Theory for Agent Behavior Analysis at Scale," is that we can use grounded theory to create an auditable trail from data right to a structured taxonomy for agent failures and successes.

Jane: This means we are moving toward methods that don't just report what happened, but actually build a framework of understanding around *why* it happened in complex scenarios.

Lu: It’s the application of this six-decade-old method to agent trajectories at scale that gives the paper its unique contribution, creating a principled way to generate behavioral insights.

Meng: I think the practical impact lies in giving developers a more systematic way to debug and understand agent behavior when they hit unexpected walls in deployment.

Lalam: This work helps solidify our ability to model agent culture by providing a rigorous, data-driven taxonomy of interaction patterns that we can use for refinement.

Tom: That’s it for today on this paper; "Using Grounded Theory for Agent Behavior Analysis at Scale." We’ll keep an eye out for what these authors do next.

Zhuoran Lu, Yangyang Yu, Zhuoyan Li, Yibo Meng, Nan Jiang, Chengxi Zang, Jie Gao and Ziang Xiao

Purdue University · Stevens Institute of Technology · Cornell University · University of Texas at El Paso · Johns Hopkins University

cs.CL, cs.AI

Submitted: 2026-08-31

Updated: 2026-08-31

Comments: 33 pages. Accepted to the Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

Code: https://github.com/ZhuoranLu/Qual-Agent-Behavior-Analysis

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: ===CATEGORIES=== ["category name": "Difficulty in task management", "definition": "This captures instances where the sequence of steps or required inputs leads to procedural failure.", "member

Key concepts

AutoTraceGT
This is a multi-agent pipeline that automates grounded theory applied to agent trajectories. It iteratively performs open, axial, and theoretical coding until saturation is reached, allowing the data itself to build the taxonomy.
Grounded Theory
A qualitative research method used here to discover underlying theories from data. The process involves constant comparison against a running codebook driven by strategic sampling to generate a behavioral taxonomy.
Agent Behavior Analysis at Scale
This refers to the application of grounded theory methods to analyze the behavior of agents across large datasets. The goal is to move beyond simple description toward building a structured framework for understanding agent decision-making and failures.
Behavioral Taxonomy
A structured categorization system that is automatically generated by the method. It captures instances where sequences of steps or required inputs lead to procedural failure, providing a framework for understanding agent interactions.

Terminology

Summary

===CATEGORIES===

[

category name: Difficulty in task management,

definition: This captures instances where the sequence of steps or required inputs leads to procedural failure.,

member codes: [trajectory id: T101, code: Procedural breakdown due to conflicting instructions],

status distribution: both,

status differentiation: In failed instances, the breakdown is characterized by an inability to prioritize steps; in resolved instances, the failure is traced back to a single, unaddressed prerequisite.

,

category name: Ambiguity of required output format,

definition: This captures moments where the necessary structure or type of information cannot be definitively determined.,

member codes: [trajectory id: T102, code: Unspecified output schema],

status distribution: failed only,

status differentiation: N/A

,

category name: Overload of competing constraints,

definition: This describes the point where too many rules or limitations are presented simultaneously, causing cognitive paralysis.,

member codes: [trajectory id: T103, code: Too many conflicting directives],

status distribution: both,

status differentiation: In failed instances, the subject attempts to satisfy all constraints equally; in resolved instances, the subject successfully identifies and prioritizes the single most critical constraint.

]

===RELATIONSHIPS===

[

categories: [Difficulty in task management, Overload of competing constraints],

relationship: When a high volume of conflicting directives are present, they directly contribute to procedural breakdown by overwhelming the system's ability to sequence actions.,

evidence: T103 (Overload) leads to T101 (Procedural breakdown due to conflicting instructions)

]

===UNDERDEVELOPED===

[

trajectory id: T205,

code: Successful integration of disparate sources,

note: "This code represents a successful outcome that requires synthesizing information from multiple, previously unrelated conceptual domains. The next batch needs to show explicit mechanisms for this synthesis, rather than just the end result."

]

===AXIAL MEMO===

The core structure emerging is a tension between the volume of available rules and the clarity of the objective. While high volumes of constraints are common, success does not seem to correlate with simply listing more rules; rather, it appears to depend on the ability to identify a single, governing hierarchy among them. The ambiguity remains whether this hierarchy is internal (a cognitive process) or external (a structural feature of the task design). The next batch needs to show explicit evidence of successful hierarchy formation—a mechanism that prioritizes rules over mere accumulation.

===END===

Improvements for AI systems

===CATEGORIES===

[

category name: N/A - Missing Data,

definition: Cannot determine behavioral quality without input codes and memos.,

member codes: [trajectory id: N/A, code: No codes provided],

status distribution: unknown,

status differentiation: n/a,

thin: true,

category memo: "<A full analysis requires the codebook and associated memo passages to establish invariance and variation. Currently, no data is available to ground any conceptual category.>"

]

===RELATIONSHIPS===

[

categories: [N/A, N/A],

relationship: Cannot establish relationships without specific codes or memo evidence.,

evidence: ""

]

===UNDERDEVELOPED===

[

trajectory id: N/A,

code: N/A,

note: "<No data provided to identify underdeveloped codes. Please supply the full dataset for review.>"

===AXIAL MEMO===

<The analysis framework is ready, but the necessary source material (codebook, memo passages, and trajectory IDs) was not included in this request batch. To proceed with selective coding or axial memo generation, please provide the data that details how concepts evolved across different experimental runs (trajectories). The next batch of data must include concrete examples of system behavior and corresponding researcher memos to allow for the identification of core theoretical mechanisms.>

===END===

Abstract

Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, often unfamiliar tasks where pre-built classifiers fall short. We propose to bring grounded theory into agent trajectory analysis: a six-decade-old qualitative method from the social sciences, with a principled saturation criterion and an auditable trail from data to theory. We propose AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories. It iteratively performs open, axial, and theoretical coding until saturation, producing a behavioral taxonomy tailored to each task. Across six trajectory corpora, AutoTraceGT produces codebooks that recover 73-91 percent of the failure modes in human-annotated taxonomies and surface additional patterns that those taxonomies miss. The emergent theoretical narrative aligns with prior expert accounts. Used as a deductive feature space, the codebook outperforms zero-shot and few-shot LLM baselines on downstream failure prediction. These results suggest Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.

Sources

Related papers