SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/ying1973/SwarmBench
Terminology
Sources
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
- MARG: Multi-Agent Review Generation for Scientific Papers
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
- LLM Multi-Agent Systems: Challenges and Open Problems
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives
- MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
- Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems
- AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
- Benchmarking LLMs' Swarm intelligence
- From Context to Skills: Can Language Models Learn from Context Skillfully?
- Kimi K2.5: Visual Agentic Intelligence
- Agent Workflow Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering