Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

arXiv:2608.12440 · cs.SE, cs.AI · Submitted 2026-08-15 · Read on arXiv

Joel Abenhaim

AI Sovereign Labs

cs.SE, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 14 pages, 4 figures, 3 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated

Terminology

Summary

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no preexisting oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully.

The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication.

The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings.

The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430.

The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.

Keywords: AI coding agents; large-scale refactoring; architectural invariants; software maintenance; agentic software engineering; verification loops.

Improvements for AI systems

Improvements to AI systems:

  1. Specification-first protocol with iterative self-audit loops – Implement a structured pipeline where the AI agent first generates a formal specification, then runs multiple refinement cycles auditing that specification against the source code before implementation, followed by atomic code changes and verification cycles against the frozen spec. This separates what to do from how to do it, reducing drift and enabling systematic defect detection.

  2. Empirical convergence criterion for autonomous verification – Replace fixed iteration counts with a stopping rule based on consecutive verification passes returning zero findings. This allows the AI to know when it has reached a stable, correct state without human intervention, and can be generalized to other long-horizon tasks (e.g., code review, theorem proving, data pipeline validation).

  3. Cross-lingual consistency checking of audit logs – Enable the AI to parse and validate its own reasoning traces (e.g., French session logs) against the specification and code, using a secondary model or self-consistency checks to detect logical gaps, contradictions, or missed edge cases. This improves transparency and trustworthiness of autonomous agents.

  4. Large-scale architectural invariant dismantling via decomposition – Train AI agents to handle changes that break central invariants by first modeling the invariant's dependencies (data flow, call graph, state transitions), then planning a phased extraction (e.g., moving state out of the panel lifecycle) rather than attempting incremental edits. The agent should be able to assess infeasibility of incremental approaches and propose a rewrite-like strategy within the same codebase.

  5. Cost-aware autonomous operation with bounded resource usage – Incorporate budget and time constraints into the agent's planning (e.g., USD 2,430, 3 days) so it can prioritize verification cycles, minimize redundant audits, and allocate compute efficiently across refinement and verification phases.

  6. Defect correction during pre-execution audits – Enhance the AI's ability to find and fix defects (201 in this case) before any human runs the program, by integrating static analysis, type checking, and behavioral simulation into the audit loop. This reduces post-deployment bugs and improves reliability for production-scale changes.

What the improved AI system can do:

  • Autonomously refactor large, interdependent codebases (e.g., 700k+ lines, 3,600+ files) to remove or alter core architectural invariants, without human code review, while maintaining functional equivalence.

  • Produce and maintain a formal, machine-readable specification that is continuously cross-checked against source code, enabling traceability and reducing hidden assumptions.

  • Self-terminate its verification process based on empirical stability (two clean passes), avoiding over- or under-iteration.

  • Detect and correct its own logical errors across multiple audit passes, including subtle issues in state management, concurrency, or lifecycle transitions.

  • Operate within explicit time and cost budgets, making it feasible for real-world production use.

  • Provide auditable, multi-lingual logs that can be automatically checked for consistency, allowing third-party models or humans to validate the agent's reasoning post-hoc.

  • Handle tasks previously deemed rewrite-only by decomposing them into extraction and reattachment phases, preserving live streams and state across UI lifecycle changes.

Abstract

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication. The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings. The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430. The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.

Sources

Related papers