Safety Invariants for Agents Orchestrating Irreversible State Transitions: A Four-Dimensional Formalism Evaluated on Public Ledgers
Zhaoming Yin
cs.CR
Submitted: 2026-08-01
Comments: 29 pages, 8 figures, 7 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autonomous agents are increasingly asked to produce irreversible effects on external systems - transferring funds, writing to durable storage, actuating hardware.
Terminology
Abstract
Autonomous agents are increasingly asked to produce irreversible effects on external systems - transferring funds, writing to durable storage, actuating hardware. Existing agent frameworks (ReAct, Reflexion, MCP) optimize task success on benchmarks and give little attention to the safety of irreversible side-effects. We formalize one such setting, movement of value across public ledgers, as state transitions in a four-dimensional space indexed by (wallet, chain, address, protocol), and use that formalism to state and prove a guarantee we call execution fidelity: under a fault model admitting planner mis-mapping, ambiguous outcomes, retries, at-least-once delivery, and delegated non-human callers, a session's realized effect on the ledger is either nothing at all or exactly the transition that was rendered to the user, exactly once. The theorem deliberately does not claim that the rendered transition matches the user's intent - no runtime layer can decide that - but it confines that unbounded question to a single predicate over a finite object, which is what makes a preview a sufficient control rather than a formality. Seven safety invariants, derived from the fidelity condition rather than enumerated from experience, discharge the guarantee. Empirically, on a controlled N=60 adversarial suite the stack lifts pass rate by 74 percentage points over a naive-ReAct baseline on two write-aggressive backing models, but by only 3 points on a write-cautious one - evidence that single-model evaluations of agent safety stacks are close to unfalsifiable. The system is deployed; 108 production write operations across 8 chains back the failure taxonomy. Although the evaluation setting is public ledgers, the formalism and invariants apply to any probabilistic agent acting on irreversible external state.
Sources
- Autonomous Agents on Blockchains: Standards, Execution Models, and Trust Boundaries
- The Price of Interoperability: Exploring Cross-Chain Bridges and Their Economic Consequences
- Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- LLMs Get Lost In Multi-Turn Conversation
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- SoK: Security of Cross-chain Bridges: Attack Surfaces, Defenses, and Open Problems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs