From Agent Output to Authorized Transition
cs.SE, cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/GoogleCloudPlatform/k8s-aibom
Project page: https://open-policy-agent.github.io/gatekeeper/website/docs/operations
Terminology
Sources
- Agent-Integrated Software: Interaction Contracts and Continuous Assurance
- Trustworthy AI Posture (TAIP): A Framework for Continuous AI Assurance of Agentic Systems at Horizontal and Vertical scale
- Runtime Governance for AI Agents: Policies on Paths
- Verifiable Manifest Signing and Transparency Enforcement for Secure MCP-Based LLM Pipelines
- Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Agentic Agile-V: From Vibe Coding to Verified Engineering in Software and Hardware Development
- Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
- Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification
- DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture
- Deontic Policies for Runtime Governance of Agentic AI Systems
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
- Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- AgentSecBench: Measuring Prompt Injection, Privacy Leakage, and Tool-Use Integrity in LLM Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties