CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
cs.LG, cs.CL
Submitted: 2026-08-06
Comments: Dataset: https://huggingface.co/datasets/AweAI-Team/CalibForge. Repository: https://github.com/AweAI-Team/CalibForge
Code: https://github.com/AweAI-Team/CalibForge
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning.
Terminology
Abstract
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.
Sources
- What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
- Building Effective AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned
- Toward Autonomous Long-Horizon Engineering for ML Research
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
- TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
- Computer Environments Elicit General Agentic Intelligence in LLMs
- Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
- Composer 2 Technical Report
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Toward Scalable Terminal Task Synthesis via Skill Graphs
- Endless Terminals: Scaling RL Environments for Terminal Agents
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- Tmax: A simple recipe for terminal agents
- Dynabench: Rethinking Benchmarking in NLP
- ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
- Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks