Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
cs.AI, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/harbor-framework/terminal-bench-2
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
- Toward Scalable Terminal Task Synthesis via Skill Graphs
- Endless Terminals: Scaling RL Environments for Terminal Agents
- SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- Tmax: A simple recipe for terminal agents
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms
- Recursive Synthesis for Long-Horizon Terminal Tasks
- CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
- CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training
- LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
- ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
- On Data Engineering for Scaling LLM Terminal Capabilities
- OpenThoughts-Agent: Data Recipes for Agentic Models
- SETA: Scaling Environments for Terminal Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection