JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
cs.CL, cs.LG
Submitted: 2026-08-26
Updated: 2026-09-03
Code: https://github.com/bingreeky/JIThttps:
Project page: https://bingreeky.github.io/JIT-sitehttps://github.com/bingreeky/JIThttps://huggingface.co/JIT-Agent
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent capability is not determined by the model alone.
Terminology
Abstract
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.
Sources
- ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems
- ClawGym: A Scalable Framework for Building Effective Claw Agents
- PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
- Automated Design of Agentic Systems
- Memory in the Age of AI Agents
- HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Recursive Harness Self-Improvement
- Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer
- Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering