Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones

arXiv:2605.03788 · cs.AI, cs.NI, cs.RO · Submitted 2026-05-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones".

Jane: The paper was written by Andrea Iannoli, Lorenzo Gigli, Luca Sciullo, Angelo Trotta and Marco Di Felice from Department of Computer Science and Engineering, University of Bologna.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are starting today with a paper that sounds like something from a sci-fi film, titled "Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones."

Jane: It does have a very cinematic feel to it, Tom, but Andrea Iannoli and his team at the University of Bologna are working on something quite practical.

Tom: They are looking at how we can move away from people having to write lines of code just to get a group of drones to move.

Jane: Exactly, because instead of programming every single tiny movement, you could just use natural language to tell the swarm what your goal is.

Lu: I think the "Web-of-Drones" concept is a beautiful way to frame this because it treats these machines as interconnected entities rather than isolated tools.

Tom: Do you mean they are essentially turning a group of drones into a single, intelligent organism?

Lu: That is a great way to look at it, since they use standardized descriptions to make sure every part of the swarm can understand and talk to every other part.

Meng: I have to wonder if that level of abstraction is actually safe when you are dealing with hardware that can crash into things.

Jane: That is a valid concern, Meng, which is why the researchers focus so much on the reasoning layer rather than just letting the AI fly wildly.

Meng: If the AI misunderstands a command like "cover this area," how do we prevent it from sending all those drones into a single collision?

Jane: They address that by using structured interfaces so the AI isn't just guessing what a drone can do or where it is.

Lu: It is like giving the AI a very specific set of rules and tools instead of just letting it wander around in its own imagination.

Meng: I will be interested to see if those rules actually hold up when things get unpredictable in the real world.

Lalam: This shift is quite profound because it changes our role from being technical operators to being high-level directors of intent.

Tom: That is a massive leap for how humans and robots interact, and it leads us right into how they actually build this system.

Summary: Jane: To understand how they manage that transition from language to action, we have to look at the architecture in "Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones."

Tom: They aren't just plugging a chatbot into a drone and hoping for the best.

Jane: No, they have built this "Agent Core" that acts as a middleman between the AI and the physical world.

Lu: I love how they use the Model Context Protocol to act as that bridge, connecting high-level thoughts to actual hardware.

Tom: Is that what allows them to use those W3C Web of Things standards they mentioned?

Lu: Yes, because by treating every drone and sensor as a "Thing" with its own description, the AI can discover what is available in real time.

Meng: So if I add a new type of moisture sensor to the swarm, the AI can actually see it and know how to use it?

Jane: That is exactly right, Meng, because the system uses a "reason-execute-monitor" loop.

Meng: Does that mean the AI is constantly checking back to see if its last command actually worked?

Jane: It does, so if a drone hasn't reached its destination, the AI sees that telemetry and has to figure out a new plan.

Lu: It is a continuous conversation between the brain and the body of the swarm.

Meng: I am still thinking about how they stop it from looping forever if something goes wrong.

Lalam: That is where those runtime guardrails come in to keep everything within safe boundaries.

Lalam: They ensure that even if the AI gets confused, it cannot execute commands that would violate basic safety protocols or mission constraints.

Tom: It sounds like they have built a very sturdy cage around a very powerful brain, which is perfect for moving into the actual results.

Improvements: Tom: Now that we know how the system is built, let's look at what happened when they actually tested "Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones."

Jane: They used a simulator called ArduPilot to test six different models on tasks like area coverage and even smart irrigation.

Tom: And it turns out that just having a smart AI isn't enough to get the job done perfectly.

Jane: That was one of the most interesting findings, especially regarding how much they benefit from specific planning tools.

Meng: I saw that in the area coverage experiments, where models without a dedicated planner really struggled to come up with a decent strategy.

Lu: It is funny because these models are so smart at talking, but they can sometimes struggle with the actual spatial geometry of a mission!

Meng: When you give them that planning tool, though, the success rates go up significantly.

Lu: Exactly, it is like giving a brilliant thinker a calculator so they can actually do the math.

Tom: But even with tools, some models had some pretty weird issues during the testing.

Jane: Right, for example, GPT was able to reach its targets but often forgot to perform the final steps like landing or disarming.

Tom: While GLM seemed to be a real standout performer in those coverage tasks.

Jane: It actually achieved a one hundred percent success rate in some of those runs, which is incredibly impressive.

Meng: I was also looking at the collision data during the formation flight tests, and it wasn't all smooth sailing.

Lu: Some models were definitely more prone to causing collisions when they had to fly in tight spaces.

Meng: It shows that we cannot just assume a larger model is always going to be a safer pilot.

Lalam: That is why the paper notes that token usage doesn't always tell you if a model is doing a good job.

Lalam: A model might use more tokens and still fail, while a more efficient one might succeed by being more precise with its reasoning.

Tom: It really proves that the way we wrap the AI in tools and safety checks is just as important as the model itself.

Conclusion: Tom: We have reached the end of our look at "Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones."

Jane: It has been fascinating to see how they have bridged that gap between human language and complex robotic coordination.

Tom: They have shown that we can move toward a future where we command swarms through intent rather than just code.

Lu: I can see this becoming a part of our everyday infrastructure, where swarms respond to our needs as naturally as we talk to each other.

Meng: I will be watching closely to see how these standardized protocols handle the chaos of real-world hardware failures and signal drops.

Lalam: This moves us toward a culture where technology is an extension of our will rather than just a tool we have to struggle to program.

Tom: Thanks for listening to this deep dive!

Jane: We will be back very soon with our next topic.

Tom: And you definitely won't want to miss it, because next time we are leaving the sky behind and heading into the world of microscopic robotics.

Department of Computer Science and Engineering, University of Bologna

cs.AI, cs.NI, cs.RO

Submitted: 2026-05-05

Updated: 2026-05-05

Comments: 15 pages, 5 figures. This paper has been accepted for presentation at the 27th IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM 2026)

Journal ref: 2026 IEEE 27th International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM), pp. 139-148, 2026

DOI: 10.1109/WoWMoM69805.2026.00027

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: This paper presents a mission-agnostic, agent-enhanced LLM framework designed for real-time UAV swarm management.

Key concepts

Agent Core
An architectural middleman that connects high-level AI reasoning to physical hardware. By using the Model Context Protocol and W3C Web of Things standards, it allows the AI to discover available drones and sensors in real time through a continuous reason-execute-monitor loop.
Intent-based Control
A method where humans act as high-level directors rather than technical operators. Instead of programming specific movements with code, users provide natural language instructions, allowing the AI to interpret goals and manage the complex coordination of an interconnected drone swarm.
Runtime Guardrails
Safety mechanisms designed to keep AI actions within secure boundaries. These prevent the system from executing commands that violate mission constraints or safety protocols, ensuring that even if the LLM makes a mistake or becomes confused, it cannot cause hardware damage or collisions.

Terminology

Summary

This paper presents a mission-agnostic, agent-enhanced LLM framework designed for real-time UAV swarm management. It addresses the challenges of heterogeneous interfaces, limited grounding, and the need for long-running closed-loop execution by allowing users to express mission objectives in natural language that the system autonomously executes through grounded, real-time interactions.

The Proposed Architecture

The framework introduces a novel architecture that moves away from static code generation at mission initialization in favor of an autonomous reason–execute–monitor cycle. By utilizing the Model Context Protocol (MCP) and W3C Web of Things (WoT) standards, the system achieves three primary innovations:

  • N1: Interoperability support through a common, hardware-independent interaction model that governs heterogeneous drone swarms.

  • N2: Real-time system state retrieval, including both swarm drones and external data sources.

  • N3: Autonomous closed-loop reasoning, where state information is incorporated into subsequent reasoning steps to enable adaptive decision-making.

System Components

The architecture is organized into three main sub-systems that facilitate bidirectional communication between the AI agent and the physical world. The W3C WoT ecosystem provides a uniform abstraction layer where drones, sensors, and services are exposed as WoT Things via machine-readable Thing Descriptions (TDs). An MCP-based WoT Gateway acts as the exclusive bridge between the agent and this ecosystem, mediating access to resources through standardized affordances.

The Agent itself is composed of several critical elements:

  • An LLM that operates at a semantic level, reasoning over abstract concepts rather than generating low-level code.

  • An Agent Core that hosts the execution loop and maintains a persistent execution context.

  • Runtime guardrail prompts, which are conditionally activated at runtime to guide reasoning and ensure safe mission execution by addressing incomplete tasks or invalid interaction sequences.

Experimental Evaluation

The researchers conducted an extensive evaluation in a hybrid real–simulated environment using the ArduPilot framework. They compared six state-of-the-art LLMs, including GPT v5.2, DeepSeek v3.2, and GLM v4.7, across four distinct swarm mission classes:

  • Area coverage (tested both with and without task-specific planning tools).

  • Formation control involving collision-aware motion.

  • Smart irrigation requiring sensor-driven decision making and communication with ground devices.

Key Findings and Insights

The results demonstrate that while current LLMs possess strong reasoning abilities, they still struggle to achieve reliable execution when operating without explicit grounding and execution support. The study highlights several critical observations regarding model performance:

  • Task-specific planning tools substantially improve robustness, though the impact varies across different LLM architectures.

  • Smaller or latency-optimized models typically show lower task performance and may require agent helper tools to reduce verbosity and bookkeeping errors.

  • There is no clear correlation between token consumption and mission success or reliability, meaning operational cost is not an indicator of execution quality.

Ultimately, the paper concludes that agent-enhanced execution and standardized device abstractions are essential to translate natural-language intent into dependable swarm behavior.

Improvements for AI systems

1. Implementation of a W3C Web of Things (WoT) Abstraction Layer via Model Context Protocol (MCP)

  • Improved AI System Capability: The system can manage heterogeneous hardware fleets (e.g., multi-vendor UAVs, diverse sensor arrays, and ground actuators) using a single, hardware-independent interaction model. By interacting with machine-readable Thing Descriptions (TDs) rather than device-specific APIs, the AI can autonomously discover, query, and actuate new resources integrated into the environment without requiring manual driver updates or code re-configuration.

2. Integration of Empirical, Failure-Driven Dynamic Guardrail Prompts

  • Improved AI System Capability: The system can autonomously correct its own reasoning trajectories during long-running missions. Instead of relying on a static system prompt, the AI detects execution anomalies—such as stalled tool calls, incomplete task termination (e.g., failing to disarm after landing), or unverified state transitions—and dynamically injects targeted corrective prompts into its context window to redirect reasoning and ensure safety-critical compliance.

3. Deployment of Hierarchical Agent Helper Tools (Semantic Wrappers)

  • Improved AI System Capability: The system can execute complex, multi-step operational sequences with high reliability even when using smaller, latency-optimized, or resource-constrained LLMs. By wrapping low-level primitives (e.g., takeoff, goto) into high-level semantic shortcuts (e.g., wait until arrived, coordinated multi drone dispatch), the AI reduces token verbosity, minimizes bookkeeping errors, and bypasses the reasoning depth limitations of edge-oriented models.

4. Transition to a Continuous Closed-Loop Reason–Execute–Monitor Cycle

  • Improved AI System Capability: The system moves from fire-and-forget mission planning to adaptive, real-time orchestration. By utilizing the MCP gateway to continuously ingest telemetry (position, battery, mode) and sensor data back into the interaction history, the AI can perform real-time state verification and adapt its subsequent reasoning steps to environmental changes or execution failures without human intervention.

Abstract

Large Language Models (LLMs) are increasingly explored as high-level reasoning engines for cyber-physical systems, yet their application to real-time UAV swarm management remains challenging due to heterogeneous interfaces, limited grounding, and the need for long-running closed-loop execution. This paper presents a mission-agnostic, agent-enhanced LLM framework for UAV swarm control, where users express mission objectives in natural language and the system autonomously executes them through grounded, real-time interactions. The proposed architecture combines an LLM-based Agent Core with a Model Context Protocol (MCP) gateway and a Web-of-Drones abstraction based on W3C Web of Things (WoT) standards. By exposing drones, sensors, and services as standardized WoT Things, the framework enables structured tool-based interaction, continuous state observation, and safe actuation without relying on code generation. We evaluate the framework using ArduPilot-based simulation across four swarm missions and six state-of-the-art LLMs. Results show that, despite strong reasoning abilities, current general-purpose LLMs still struggle to achieve reliable execution - even for simple swarm tasks - when operating without explicit grounding and execution support. Task-specific planning tools and runtime guardrails substantially improve robustness, while token consumption alone is not indicative of execution quality or reliability.

Sources

Related papers