Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
Joel Abenhaim
AI Sovereign Labs
cs.SE, cs.AI
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 14 pages, 4 figures, 3 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated
Terminology
Summary
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no preexisting oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully.
The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication.
The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings.
The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430.
The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.
Keywords: AI coding agents; large-scale refactoring; architectural invariants; software maintenance; agentic software engineering; verification loops.
Improvements for AI systems
Improvements to AI systems:
-
Specification-first protocol with iterative self-audit loops – Implement a structured pipeline where the AI agent first generates a formal specification, then runs multiple refinement cycles auditing that specification against the source code before implementation, followed by atomic code changes and verification cycles against the frozen spec. This separates
what to do
fromhow to do it,
reducing drift and enabling systematic defect detection. -
Empirical convergence criterion for autonomous verification – Replace fixed iteration counts with a stopping rule based on consecutive verification passes returning zero findings. This allows the AI to know when it has reached a stable, correct state without human intervention, and can be generalized to other long-horizon tasks (e.g., code review, theorem proving, data pipeline validation).
-
Cross-lingual consistency checking of audit logs – Enable the AI to parse and validate its own reasoning traces (e.g., French session logs) against the specification and code, using a secondary model or self-consistency checks to detect logical gaps, contradictions, or missed edge cases. This improves transparency and trustworthiness of autonomous agents.
-
Large-scale architectural invariant dismantling via decomposition – Train AI agents to handle changes that break central invariants by first modeling the invariant's dependencies (data flow, call graph, state transitions), then planning a phased extraction (e.g., moving state out of the panel lifecycle) rather than attempting incremental edits. The agent should be able to assess infeasibility of incremental approaches and propose a rewrite-like strategy within the same codebase.
-
Cost-aware autonomous operation with bounded resource usage – Incorporate budget and time constraints into the agent's planning (e.g., USD 2,430, 3 days) so it can prioritize verification cycles, minimize redundant audits, and allocate compute efficiently across refinement and verification phases.
-
Defect correction during pre-execution audits – Enhance the AI's ability to find and fix defects (201 in this case) before any human runs the program, by integrating static analysis, type checking, and behavioral simulation into the audit loop. This reduces post-deployment bugs and improves reliability for production-scale changes.
What the improved AI system can do:
-
Autonomously refactor large, interdependent codebases (e.g., 700k+ lines, 3,600+ files) to remove or alter core architectural invariants, without human code review, while maintaining functional equivalence.
-
Produce and maintain a formal, machine-readable specification that is continuously cross-checked against source code, enabling traceability and reducing hidden assumptions.
-
Self-terminate its verification process based on empirical stability (two clean passes), avoiding over- or under-iteration.
-
Detect and correct its own logical errors across multiple audit passes, including subtle issues in state management, concurrency, or lifecycle transitions.
-
Operate within explicit time and cost budgets, making it feasible for real-world production use.
-
Provide auditable, multi-lingual logs that can be automatically checked for consistency, allowing third-party models or humans to validate the agent's reasoning post-hoc.
-
Handle tasks previously deemed
rewrite-only
by decomposing them into extraction and reattachment phases, preserving live streams and state across UI lifecycle changes.
Abstract
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication. The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings. The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430. The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.
Sources
- The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Large Language Models Cannot Self-Correct Reasoning Yet
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties