Mechanism Design for Alignment and Control
econ.TH, cs.AI, cs.GT
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Measuring Progress on Scalable Oversight for Large Language Models
- Supervising strong learners by amplifying weak experts
- Emergent Alignment via Competition
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Robust Trust
- Friend or Foe: Delegating to an AI Whose Alignment is Unknown
- Screening for Choice Sets
- Alignment faking in large language models
- Multi-Agent Risks from Advanced AI
- Radical AI Interpretability
- The Elicitation Game: Evaluating Capability Elicitation Techniques
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Risks from Learned Optimization in Advanced Machine Learning Systems
- Exploration Hacking: Can LLMs Learn to Resist RL Training?
- Risk thresholds for frontier AI
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Realistic honeypot evaluations for scheming propensity