Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
cs.SE, cs.AI
Submitted: 2026-08-22
Updated: 2026-09-20
Code: https://github.com/Parker-Fawcett/rebuild-dossier
Terminology
Sources
- AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
- SWE-Bench+: Enhanced Coding Benchmark for LLMs
- Concrete Problems in AI Safety
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties