Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/openai/codex
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Towards Iterative End-to-End Software Development: A Feature-Driven Multi-Agent Framework
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
- MemGPT: Towards LLMs as Operating Systems
- ProgramBench: Can Language Models Rebuild Programs From Scratch?
- Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment
- Self-Harness: Harnesses That Improve Themselves
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection