RAT: RunAnyThing via Fully Automated Environment Configuration
cs.SE, cs.AI
Submitted: 2026-04-25
Updated: 2026-09-19
Code: https://github.com/bndr/pipreqs
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automating repository-level software engineering tasks is a foundational challenge for autonomous code agents, largely due to the difficulty of configuring executable environments.
Terminology
Abstract
Automating repository-level software engineering tasks is a foundational challenge for autonomous code agents, largely due to the difficulty of configuring executable environments. However, manual configuration remains a labor-intensive bottleneck, necessitating a transition toward fully automated environment configuration. Existing approaches often rely on pre-defined artifacts or are restricted to specific programming languages, limiting their applicability to diverse real-world repositories. In this paper, we first propose RAT (RunAnyThing), a modular and extensible agent framework for fully automated configuration across programming languages on arbitrary repositories. RAT adopts a multi-stage pipeline that integrates language-aware abstraction, image initialization, specialized configuration toolset, and robust sandbox. Furthermore, to enable rigorous evaluation, we propose RATBench, a benchmark reflects the comprehensive coverage of real-world repositories. Extensive experiments demonstrate that RAT achieves state-of-the-art performance, improving Environment Setup Success Rate (ESSR) by an average of 36.1% over strong baselines.
Sources
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- EnvBench: A Benchmark for Automated Environment Setup
- Repo2Run: Automated Building Executable Environment for Code Repository at Scale
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
- StarCoder: may the source be with you!
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
- GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
- Repoformer: Selective Retrieval for Repository-Level Code Completion
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- Qwen3 Technical Report
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties