Arbiter: Detecting Interference in LLM Agent System Prompts

arXiv:2603.08993 · cs.SE, cs.AI, cs.CR, cs.PL · Submitted 2026-03-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Arbiter: Detecting Interference in LLM Agent System Prompts".

Tom: The gist The agent that resolves the conflict cannot be the agent that detects it.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, Arbiter is basically a way to find problems in those big system instructions we give AI agents before we even run them, Jane.

Jane: Right, it’s about checking the blueprint—the prompt itself—for internal contradictions that might cause the agent to just guess its way out of a mess instead of actually solving the problem Tom.

Lu: It’s looking at how these prompts are built, like monolithic ones that get messy at the boundaries or modular ones where pieces don't connect right Jane.

Meng: So you’re saying we need formal rules to check if one part of the prompt is accidentally contradicting another part before we deploy it to real work?

Tom: Exactly. It uses these formal rules, like checking for scope overlap or priority conflicts, and then it uses these undirected scouring tests where different AIs look at the same prompt with vague instructions Jane.

Jane: That’s smart because it helps you see things a single model might miss, bringing in different analytical angles about what's actually problematic Lu.

Lu: And the big finding is that those prompt structures directly map to the type of bug you get, so if your prompt is flat, you expect capability trade-offs and if it’s modular, you see structural issues where parts don't communicate properly Tom.

Meng: From an engineering standpoint, that means we need to treat our system prompts like software architectures where the way you assemble the pieces determines how stable the whole thing is Jane.

Tom: And the cost of doing this deep check was surprisingly low, less than a quarter of a dollar across all three big vendor prompts Lu, which shows it’s something any developer can actually do easily now.

Jane: It really changes how we think about prompt engineering from just 'making it work' to making sure the instructions themselves are structurally sound Tom.

Lu: And they found that because different models have different training, they each end up inventing their own way to categorize these interference patterns, which shows how model bias influences our detection Jane.

Meng: It makes me think about building testing infrastructure for prompts, like linters or specific checks to monitor how the structure of the prompt changes over time Tom.

Tom: Yeah, and they’re suggesting we need a way to track those structural changes precisely with tools that look at the underlying code, not just line-by-line text edits Jane.

Jane: And they want a scale for severity so you know if a finding is something you need to worry about operationally or if it's just interesting data points Lu.

Lu: So, the main thing is that there’s a clear link between how you write these instructions and the kinds of bugs that show up in the AI agent's behavior Tom.

Tom: It’s about moving away from just hoping the AI handles conflicts well and toward engineering those instructions for reliability Jane.

Jane: It opens up a lot of possibilities for how we design more robust agents, especially if you can start tracking prompt evolution structurally rather than just reading it as text Lu.

Lu: And they point out that because different models have different training, they each end up inventing their own way to categorize these interference patterns Jane.

Meng: It makes me think about building testing infrastructure for prompts, like linters or specific checks to monitor how the structure of the prompt changes over time Tom.

Tom: Yeah, and they’re suggesting we need a way to track those structural changes precisely with tools that look at the underlying code, not just line-by-line text edits Jane.

Jane: And they want a scale for severity so you know if a finding is something you need to worry about operationally or if it's just interesting data points Lu.

Lu: It’s about making the evaluation process itself more rigorous, ensuring that when you look at those one hundred fifty-two findings in Arbiter, you aren't just looking at a list of things, but weighted problems Jane <ref:2603.08993#pg1>.

Tom: So to wrap up this part of the talk on Arbiter: Detecting Interference in LLM Agent System Prompts, the core idea is that prompt architecture dictates what kind of failure mode you’re likely to see, and using a combination of formal rules and multi-model scouring helps uncover those different failure classes Jane.

Jane: It’s about moving away from just hoping the AI handles conflicts well and toward engineering those instructions for reliability Lu.

Lu: It opens up a lot of possibilities for how we design more robust agents, especially if you can start tracking prompt evolution structurally rather than just reading it as text Meng.

Meng: I agree, having better tools to monitor prompt evolution means fewer surprises when an agent starts acting weird in a complex workflow Jane.

The paper's summary: Tom: So, Arbiter isn't just about finding problems in prompts; the authors are suggesting how we should actually improve these systems for better long-term reliability Jane.

Jane: Right, they’re saying we need to move beyond just checking if a prompt works today and start building it with checks built in from the beginning Tom.

Lu: They propose this version control thing, using something like an Abstract Syntax Tree to track exactly what changed structurally in the prompt, not just line by line text edits Jane.

Meng: That sounds really practical for engineering; you could actually see a structural diff of the AI's instructions as you iterate on them Lu.

Tom: And that’s huge because it lets us classify those changes—added, removed, modified—based on how they affect the structure itself, not just what words were swapped around Jane.

Jane: It means we could pinpoint exactly where a design seam is getting weaker before the agent even runs in production Tom.

Lu: Plus, they’re pushing for a clear way to score findings with both severity and confidence so you can actually weigh the issues correctly Lu.

Meng: If we get that scoring system, it helps us prioritize which structural flaws we need to fix first when we have too much to deal with Jane.

Tom: And they’re emphasizing this need for formal evaluation rules upfront, so the AI doesn't just use its judgment to smooth over a conflict between two instructions Lu.

Jane: That makes sense because if the prompt has an internal contradiction, you don't want the AI to just guess which instruction wins without a hard rule Tom.

Lu: It’s about forcing explicit checks for things like mandate-prohibition conflicts right at the design stage Jane.

Meng: So, it’s not just about finding errors after they happen; it’s about building guardrails into the prompt creation process Lu.

Tom: Exactly, and this whole approach makes the cost of analysis really low, which means we can do this kind of deep structural check across many different AI systems easily Jane.

Jane: It puts a lot more responsibility on us to treat these instructions like actual software code that needs rigorous testing Tom.

Lu: And Lalam is seeing this as a cultural shift; if we can engineer these guardrails into the instructions, it improves how you interact with and trust the AI in your daily work Lu.

Meng: I agree, having better tools to monitor prompt evolution means fewer surprises when an agent starts acting weird in a complex workflow Jane.

Tom: So, the big picture is that Arbiter gives us a blueprint for how to engineer system prompts more carefully before we let them run unsupervised Lu.

The paper's improvements: Tom: So we’re wrapping up on Arbiter: Detecting Interference in LLM Agent System Prompts, which is this framework for detecting interference in LLM agent system prompts Jane.

Jane: Right, essentially it’s showing us that the way you write those instructions directly dictates the kind of bugs we see later when the AI starts working Tom.

Lu: It really highlights that prompt architecture matters just like software architecture does for conventional code, Lu.

Meng: It means we need to think about these instructions structurally before we let them run in a complex environment Jane.

Lalam: From my side, this suggests that better prompt engineering will lead to more reliable AI interactions overall, improving how you use these tools in our culture Lalam.

Tom: The numbers show that the cost of doing this thorough cross-vendor check is really low, less than a quarter of a dollar Lu.

Jane: And they’re pushing for things like structural analysis tools using Abstract Syntax Trees to track prompt changes precisely Tom.

Lu: It opens up possibilities for monitoring prompt evolution in real time, which is pretty exciting from a research standpoint Lu.

Meng: I think being able to monitor those structural shifts is what makes this useful for building production systems Jane.

Tom: We've seen how monolithic prompts cause growth-level bugs and modular ones cause design-level issues with composition seams Lu.

Jane: So the big picture is that we need more formal ways to test these instructions before we let them run unsupervised Tom.

Lu: It sets a new standard for prompt engineering by focusing on the structure itself Jane.

Tom: The data is clear, and nobody’s really checking this kind of structural detail yet Lu.

Jane: We’ve got a lot of ground to cover, but this shows us where the next big engineering challenges are with these models Tom.

Lu: Next up, we want to look at how multimodal models are actually handling physical fields, like that PhysFieldBench paper Jane.

Conclusion: Tom: So we've been looking at Arbiter today, which is this framework for detecting interference in LLM agent system prompts Jane.

Jane: Right, essentially it’s showing us that the way you write those instructions directly dictates the kind of bugs we see later when the AI starts working Tom.

Lu: It really highlights that prompt architecture matters just like software architecture does for conventional code, Lu.

Meng: It means we need to think about these instructions structurally before we let them run in a complex environment Jane.

Lalam: From my side, this suggests that better prompt engineering will lead to more reliable AI interactions overall, improving how you use these tools in our culture Lalam.

Tom: The numbers show that the cost of doing this thorough cross-vendor check is really low, less than a quarter of a dollar Lu.

Jane: And they’re pushing for things like structural analysis tools using Abstract Syntax Trees to track prompt changes precisely Tom.

Lu: It opens up possibilities for monitoring prompt evolution in real time, which is pretty exciting from a research standpoint Lu.

Meng: I think being able to monitor those structural shifts is what makes this useful for building production systems Jane.

Tom: We've seen how monolithic prompts cause growth-level bugs and modular ones cause design-level issues with composition seams Lu.

Jane: So the big picture is that we need more formal ways to test these instructions before we let them run unsupervised Tom.

Lu: It sets a new standard for prompt engineering by focusing on the structure itself Jane.

Tom: The data is clear, and nobody’s really checking this kind of structural detail yet Lu.

Jane: We’ve got a lot of ground to cover, but this shows us where the next big engineering challenges are with these models Tom.

Lu: Next up, we want to look at how multimodal models are actually handling physical fields, like that PhysFieldBench paper Jane.

University of British Columbia · Georgia Institute of Technology

cs.SE, cs.AI, cs.CR, cs.PL

Submitted: 2026-03-09

Updated: 2026-10-07

Importance score: 90/100

The gist: The gist The agent that resolves the conflict cannot be the agent that detects it.

Key concepts

Monolithic Prompts
These are large, single system prompts. They are compared to monolithic software applications where adding features independently leads to contradictions at subsystem boundaries. This architecture is linked to 'growth-level bugs' that appear when different parts of the prompt conflict.
Modular Prompts
These prompts are built from smaller, composed functions or blocks. They are analogous to modular software design where contracts between modules might be missing. This structure is associated with 'design-level bugs' occurring at the seams where these separate components connect.
Undirected Scouring
This phase involves sending a system prompt to multiple LLMs with vague instructions, asking them to note interesting patterns and rate severity. It helps uncover unexpected issues by leveraging different models' analytical biases, revealing what might be 'concerning' or 'alarming'.

Terminology

Summary

The gist The agent that resolves the conflict cannot be the agent that detects it. <ref:2603.08993#pg14>

System Prompt Analysis and Failure Modes

The paper presents Arbiter, a framework combining formal evaluation rules with multi-model LLM scouring to detect interference patterns in system prompts. <ref:2603.08993#pg15> Applied to three major coding agent system prompts—Claude Code (Anthropic), Codex CLI (OpenAI), and Gemini CLI (Google)—the analysis identified 152 findings across the undirected scouring phase and 21 hand-labeled interference patterns in directed analysis of one vendor. <ref:2603.08993#pg16> The findings organize into a taxonomy correlated with prompt architecture: monolithic prompts produce growth-level bugs at subsystem boundaries, flat prompts trade capability for consistency, and modular prompts produce designlevel bugs at composition seams. <ref:2603.08993#pg17> This taxonomy is grounded in the analogy between system prompt architectures and software architectures, where Monolithic applications accumulate contradictions at subsystem boundaries as teams add features independently <ref:2603.08993#pg18>.

Evaluation Methodology

Arbiter employs two complementary evaluation phases: directed evaluation and undirected scouring. <ref:2603.08993#pg19> Directed evaluation decomposes a system prompt into classified blocks and evaluates block pairs against formal interference rules, which includes Mandate-prohibition conflict, Scope overlap, and Priority ambiguity. The undirected scouring phase sends the prompt to multiple LLMs with deliberately vague instructions, asking them to note what they find interesting and rate severity on a four-level epistemic scale: curious (pattern noticed), notable (worth investigating), concerning (likely problematic), alarming (structurally guaranteed to cause failures).

Architectural Correlations and Findings

The central finding is that prompt architecture strongly correlates with observed failure mode class. A monolithic prompt, like Claude Code’s 1,490-line document, produces growth-level bugs at subsystem boundaries, such as the critical contradictions between a general-purpose subsystem and specific workflow subsystems. Conversely, a modular prompt, exemplified by Gemini CLI’s composition from TypeScript render functions, exhibits design-level bugs at composition seams, such as structural data loss during history compression where the contract between modules is never written.

Multi-Model Complementarity

The multi-model evaluation phase reveals that different models bring different analytical biases rooted in their training data and architecture. This complementarity is demonstrated by the fact that The category explosion in Claude Code—107 unique categories for 116 findings—quantifies this: each model invents its own taxonomy. Findings are clustered into meta-categories, showing that models are complementary, not redundant.

Economic Validation and Conclusion

The total cost of cross-vendor analysis was remarkably low at a total of 27 cents USD, which is less than three minutes of minimum wage labor. This demonstrates that system prompt analysis at this level of thoroughness is accessible to any developer with API access. The paper concludes that The tools exist. The data is clear. Nobody is checking.

Limitations and Implications

The analysis is limited to static analysis, as it examines what the prompt says, not runtime behavior. However, the structural conditions identified suggest that system prompts need the same engineering infrastructure that conventional software has: linters for internal consistency, tests for behavioral contracts. The paper highlights responsible disclosure regarding findings like the Gemini CLI’s save memory data loss, which is a structural guarantee that saved preferences are deleted during history compression. The overall cost per finding across all three analyses is calculated at less than three minutes of human labor.

The deterministic reproduction workflow and artifact manifest checks are documented in ARTIFACT.md in the repository. The total API cost across all three analyses was less than three minutes of human labor. The paper concludes that The tools exist. The data is clear. Nobody is checking. The total cost of comprehensive cross-vendor analysis was twenty-seven cents—less than three minutes of minimum wage labor, less than a single finding from a human security audit. The tools exist. The data is clear.

Improvements for AI systems

  1. System prompts should be subjected to formal evaluation rules to detect internal contradictions before deployment. This prevents agents from silently resolving conflicts through heuristics, as an LLM executing a contradictory system prompt will smooth over inconsistencies through “judgment”—the same mechanism that makes LLMs useful makes them unreliable as their own auditors.

  2. Implement multi-model scouring for undirected analysis to discover vulnerability classes missed by single-model evaluations. This approach addresses the gap where Directed rules find exhaustive enumerations but Undirected scouring finds what’s outside the search frame, uncovering issues like security architecture gaps (system-reminder trust as injection surface).

  3. Adopt a prompt architecture that favors modularity to mitigate design-level bugs at composition seams. This is achieved by recognizing that for modular prompts, the bugs exist exclusively in the gaps between modules—contracts that were never specified because each module was designed independently.

  4. Develop version control and structural analysis tools based on an Abstract Syntax Tree (AST) to monitor prompt evolution precisely. This allows for AST diffing to classify every node as added, removed, modified (same structural hash, different content), or moved (same content hash, different structural hash) without relying on line-level text diffing.

  5. Establish a comprehensive severity and epistemic confidence scale for all findings to weight results by source reliability. This allows consumers to differentiate between findings where A finding can be epistemically “curious” but operationally critical, or epistemically “alarming” but operationally irrelevant.

Sources

Related papers