Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
cs.AI, cs.CR, cs.ET
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 19 pages
Code: https://github.com/NVIDIA/nvtrust
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor.
Terminology
Abstract
Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation of assurance,is still applied to them as a binary. We argue that it must be graded along three orthogonal axes: principal independence (who controls the auditor), substrate independence (an auditor sharing the auditee's foundation-model family, toolchain or guardrails fails with it) and evidence independence (whether evidence is attestable rather than self-reported). Each axis has precedent; the contribution is to grade all three on a single audit, aggregate them by the weakest link, and apply the same rubric when the auditor is itself an agent. We give the model a formal basis by transplanting the beta-factor model of common-cause failure from reliability engineering, a seven-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement-controls agent audited at three grades, and a Monte Carlo study of the model in which a conventional internal audit of an agent-a real audit team, a second agent, provider logsp-surfaces 5.9% of the faults it could in principle see and none at all in half the fault classes. We map the triple to the EU AI Act as amended, ISO/IEC 42006, UK public-sector risk-management guidance and audit-regulator practice.
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Ctrl-Z: Controlling AI Agents via Resampling
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
- Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
- IDs for AI Systems
- Infrastructure for AI Agents
- When can we trust untrusted monitoring? A safety case sketch across collusion strategies
- Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software
- Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach
- AI Control: Improving Safety Despite Intentional Subversion
- Subversion via Focal Points: Investigating Collusion in LLM Monitoring
- AI Agents That Matter
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
- Governing AI Agents
- Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing
- Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
- Who Should Run Advanced AI Evaluations -- AISIs?
- Self-Preference Bias in LLM-as-a-Judge
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Governing Dynamic Capabilities: Cryptographic Binding and Reproducibility Verification for AI Agent Tool Use
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection