AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering
cs.SE, cs.AI, cs.MA
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/KaituoZhang/AsynCodeBench
Terminology
Sources
- Why Do Multi-Agent LLM Systems Fail?
- AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- Effective Strategies for Asynchronous Software Engineering Agents
- TeamBench: Evaluating Agent Coordination under Enforced Role Separation
- AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
- Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties