BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
cs.SE, cs.AI, cs.PL
Submitted: 2025-09-27
Updated: 2026-09-16
Comments: Accepted at TMLR
Code: https://github.com/huzecong/ghcc
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Language Models are Few-Shot Learners
- Why Do Multi-Agent LLM Systems Fail?
- CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
- Self-Organized Agents: A LLM Multi-Agent Framework toward Ultra Large-Scale Code Generation and Optimization
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Large Language Models are Zero-Shot Reasoners
- DIRE: A Neural Approach to Decompiled Identifier Naming
- Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
- CoverUp: Effective High Coverage Test Generation for Python
- Reflexion: Language Agents with Verbal Reinforcement Learning
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
- Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
- LLaMA: Open and Efficient Foundation Language Models
- Agent Workflow Memory
- ReAct: Synergizing Reasoning and Acting in Language Models
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties