Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
cs.SE, cs.AI
Submitted: 2026-08-08
Updated: 2026-08-31
Code: https://github.com/razzant/ouroboros
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work.
Terminology
Abstract
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.
Sources
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- Automated Design of Agentic Systems
- AI safety via debate
- Constitutional AI: Harmlessness from AI Feedback
- Self-Refine: Iterative Refinement with Self-Feedback
- Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
- Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception
- GAIA: a benchmark for General AI Assistants
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- A Self-Improving Coding Agent
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
- G\"odel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement
- The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties