Algorithmic Unverifiability of Safety for Fixed and Recursively Self-Improving Systems
cs.LO, cs.AI, cs.CC, cs.CL
Submitted: 2026-06-26
Updated: 2026-09-23
Terminology
Sources
- No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees
- Concrete Problems in AI Safety
- Supervising strong learners by amplifying weak experts
- AI safety via debate
- Scalable agent alignment via reward modeling: a research direction
Related papers
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification
- A programming language combining quantum and classical control
- Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows
- Encoder-Decoder Transformers: Logical Characterizations and Periodicity
- Ultraconstructive Model Theory via Bounded Adversarial Finite Structures