FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration
cs.SE, cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/codefuse-ai/codef
Terminology
Sources
- AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
- EnvBench: A Benchmark for Automated Environment Setup
- Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering
- Repo2Run: Automated Building Executable Environment for Code Repository at Scale
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- RAT: RunAnyThing via Fully Automated Environment Configuration
- ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
- Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation
- PIPer: On-Device Environment Setup via Online Reinforcement Learning
- AgentBench: Evaluating LLMs as Agents
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
- ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- NetArena: Dynamic Benchmarks for AI Agents in Network Automation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties