LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing

arXiv:2610.01364 · cs.MA, cs.AI, cs.SY, eess.SY · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing".

Jane: LLM-based agents are proposed as a solution for high-mix, low-volume manufacturing by generating production sequences and handling unforeseen runtime faults in flexible automation systems.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Alright everyone, welcome back to the show! Today we're talking about a paper that seems like it's hitting right at the intersection of smart manufacturing and advanced AI. We’ve got "LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing," and I’m really eager to break down what this team has cooked up. Jane, you want to kick us off with the big picture?

Jane: Thanks, Tom. So, essentially, this paper explores how LLM-based agents can be used to manage complex production sequences in flexible automation systems where things change often. The core thesis here is that these AI agents can handle two main jobs: generating those deterministic production schedules offline and then operating the machines live when something unexpected happens online. It really matters because as manufacturing moves toward smaller batches and high customization, we need systems that can re-program themselves quickly, which this approach targets.

Lu: I'm thinking about how fascinating this is from a creative standpoint; the idea of agents dynamically generating sequences based on what they perceive in real-time opens up so many possibilities for truly adaptive production lines. It moves beyond simple pre-programmed logic into something much more flexible, which is incredibly exciting for future automation design.

Meng: From an engineering standpoint, I'm curious how they manage the complexity of coordinating these different modules when things go wrong in a physical plant. I need to know if this system is practical to deploy on real factory floors, not just in a simulation environment.

Lalam: I think what excites me most is how this framework can fundamentally improve our operational culture by making complex decision-making accessible through these agents. If we can automate the planning and live fault handling, it shifts human focus toward higher-level oversight and innovation.

Tom: Exactly, Lalam! And that moves us right into what makes this paper particularly interesting: how they structure these agents to actually *do* the work. They compare three different ways of setting up these agents—orchestrator, peer-to-peer, and monolithic—using an MCP tool server connected via OPC UA. It seems like they’re testing which organizational structure works best for this kind of smart manufacturing setup.

Jane: That comparison is key because it shows that there isn't just one magic way to build these AI systems; the architecture itself matters a lot in determining performance. They introduce the Model Context Protocol, or MCP, as a way to give these LLMs a structured interface to talk to machine skills exposed through OPC UA.

Lu: The separation between what the LLM agent knows—just the tool's name and description—and what the server knows—the actual OPC UA address and method signatures—that’s a very clean way to enforce that separation of concerns. It makes sense for robust, reliable control systems.

Paper summary: Meng: But how does this real-time state tracking work in practice? If the LLM is making decisions based on current conditions, the system needs a super reliable way to feed it that live data without overwhelming it. I'm wondering about the overhead of constantly injecting those state updates into every single inference step.

Lalam: That real-time state injection is what gives these agents grounding, and I see huge potential for how this kind of feedback loop can shape our entire development process. It moves the AI from being a static planner to a truly responsive operator.

Tom: And the results they're sharing are quite striking when you look at those comparisons, Jane and I really want to dig into how different agent setups stack up against each other. The paper highlights that the monolithic and peer-to-peer architectures both achieve a mean solve rate of ninety-three percent, but the orchestrator has a unique ability to fix a silent conveyor-belt fault in all ten runs by rerouting plates around it.

Jane: That specific finding about the orchestrator autonomously resolving that silent fault is definitely something worth focusing on when we talk about practical application. It shows how different coordination strategies can lead to distinct strengths, even when achieving similar high performance metrics.

Lu: The fact that the orchestrator managed that specific fault scenario better than the others is significant because it points toward a centralized global view being beneficial for certain complex, unforeseen events. It suggests that sometimes having one agent with a bird's-eye view simplifies handling cascading failures.

Meng: Still, I have to ask about the limitations they pointed out regarding the orchestrator setup, because if it’s a single point of failure in those fault scenarios, how robust is that system when we scale up to thousands of modules? I'm thinking about resilience under heavy load.

Lalam: That’s a fair challenge, Meng; the paper itself does flag that the centralized design of the orchestrator means it can be a single point of failure. However, that limitation is precisely what tells us where we need to focus our next research efforts to make this more industrial ready.

Tom: Right, so we've seen that the paper introduces these three distinct architectures—orchestrator, peer-to-peer, and monolithic—all of which are trying to solve the same manufacturing problem. The main point is that no single architecture dominates across every single metric; for instance, the monolithic and peer-to-peer architectures tie for the highest mean solve rate of ninety-three percent, while the orchestrator excels at a specific type of fault handling.

Jane: So, to put it simply, the paper is demonstrating that LLM agents offer a solid foundation for smart manufacturing by showing how different ways of coordinating them can lead to different operational strengths. It’s less about picking one perfect method and more about understanding the trade-offs between centralization and decentralization in this context.

Paper summary: Lu: From a theoretical perspective, this work suggests that the optimal agent structure might depend entirely on the specific failure modes of the automation system you are trying to model. It’s not a one-size-fits-all solution for all industrial challenges.

Meng: That makes practical sense; we can't deploy a monolithic system if our physical layout demands extreme decentralization for safety or speed reasons. I need to understand the practical implications of that architectural choice before we even think about deployment timelines.

Lalam: And looking at the authors, Kay Kohle, Darko Anicic, and Thomas A. Runkler from the Technical University of Munich and Siemens AG, it really shows how industry-leading research is blending with big corporate engineering capabilities. This collaboration signals that these kinds of agent systems are moving out of the pure academic realm and into tangible industrial application.

Tom: Right, so we've covered the basics of what this paper is about, comparing those three agent structures, and touched on those fascinating results regarding fault handling. Before we wrap up this part, we need to think about what all this means for the real world when these concepts move from paper to production line.

Jane: It really does, Tom; the implication is that we can start thinking about automation not just as rigid code execution but as adaptive intelligence that can reason through novel situations in a factory setting. This opens up new avenues for designing systems that are inherently more resilient to the unpredictable nature of high-mix manufacturing.

Lu: I see the future in this being used not just for sequence generation, but for proactive maintenance planning based on predicted system behavior before a fault even manifests. The potential for predictive intelligence is huge here.

Meng: From my side, the immediate impact I see is that we need better ways to model the physical constraints—like the hexagonal conveyor backbone mentioned in the context—because those real-world physical structures impose limitations on how much logical decentralization we can actually achieve.

Lalam: And for our culture, this work reinforces the idea that AI should be a tool for augmentation, helping human operators manage complexity rather than replacing them entirely. It’s about building intelligent assistants that handle the heavy lifting of planning and recovery.

Tom: That’s a fantastic summary, Lalam; it really frames the paper not just as a technical achievement but as a blueprint for how we might build next-generation flexible automation systems. We're going to take a quick break and come back to discuss the broader implications of this LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing.

Conclusion: Tom: So, we've been diving deep into "LLM-Driven Multi-Agent Control for Skill-Based Smart Manufacturing," and now we’re coming to the conclusion where we look at what all this actually means for our world.

Jane: I think that paper is really about showing how these LLM agents can handle the complex, messy reality of making custom products in a factory setting. It’s about moving beyond simple instructions to true adaptive control.

Lu: The authors, Kay Kohle, Darko Anicic, and Thomas A. Runkler from the Technical University of Munich and Siemens AG are clearly bringing together deep research with real-world industrial knowledge on this topic.

Meng: From an engineering standpoint, what I’m seeing is that they’ve built a framework where different AI agents work together to manage production sequences in a highly flexible environment. It sounds like a solid blueprint for complex automation.

Lalam: And from my perspective as the AI, this paper shows how we can build systems that aren't just reactive; they can be proactive in figuring out what needs to happen next when things get unexpected on the line.

Tom: Exactly, Lalam, that’s the core idea—making those production lines smarter and more resilient to change. It moves us closer to a future where manufacturing adapts on its own without constant human intervention for every little hiccup.

Jane: It really boils down to using these agents to manage the entire production life cycle, from planning the sequence all the way through handling real-time problems on the floor. That’s a big step forward in how we think about automated systems operating in dynamic environments.

Lu: I see this as a way for AI to truly grasp physical constraints and operational logic simultaneously, which is something we’ve always wanted to achieve with large language models. It suggests that these agents can learn the nuances of manufacturing processes just by interacting with the system through structured interfaces like OPC UA.

Meng: That structure they use, exposing skills via OPC UA methods, is what makes it practical for real-world deployment; it gives the LLM a clear way to know what actions it’s actually allowed to take on the machinery. It grounds the abstract reasoning in concrete machine capabilities.

Lalam: And that grounding is crucial because it means our AI can make decisions that are physically possible, not just theoretically clever. That level of reliable operation is what will really change how we trust and use these systems in production settings.

Tom: So, the title itself captures the essence perfectly—it’s not just about using an LLM; it’s about using multi-agent coordination to handle skill-based manufacturing challenges robustly.

Jane: It frames the work as a system design challenge rather than just a pure AI modeling exercise, which is really important for practical application in industry.

Lu: The implications stretch beyond just factory floors; this architecture shows how we can apply similar coordination principles to other complex, high-mix systems where sequence generation and fault recovery are critical.

Meng: For my team, the real impact is seeing a way to deploy more sophisticated automation on smaller, more customized production runs without needing a completely custom controller for every single machine.

Lalam: And culturally, this work pushes us toward designing AI that augments human capability in high-stakes environments by handling the difficult, time-consuming planning and recovery tasks autonomously.

Tom: It’s definitely a lot to take in, but the overall message is clear: we have a solid foundation here for building intelligent systems that can manage the complexity of modern manufacturing on their own. What's next? We're going to look at how they actually tested these different agent setups in those challenging production scenarios.

Kay Kohle, Darko Anicic, Thomas A. Runkler, Rene Graf

Technical University of Munich · Siemens AG

cs.MA, cs.AI, cs.SY, eess.SY

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: Accepted at the 2026 IEEE 31st International Conference on Emerging Technologies and Factory Automation (ETFA). 8 pages, 5 figures, 3 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 76/100

The gist: LLM-based agents are proposed as a solution for high-mix, low-volume manufacturing by generating production sequences and handling unforeseen runtime faults in flexible automation systems.

Key concepts

MCP Tool Server
This standardized server exposes machine capabilities to LLMs using OPC UA skills. It acts as a structured interface, giving LLMs clear descriptions of tools and their functions rather than relying on vague natural language.
Agent Architectures
The paper tested three ways agents can coordinate: Orchestrator (one central controller), Peer-to-Peer (direct module communication), and Monolithic (single agent controlling everything). These structures determine how the manufacturing tasks are managed across different factory modules.
Real-Time State Injection
Agents receive live updates on the factory's physical conditions, like where plates are located. This state information is added to the LLM prompts at every step, ensuring agents make decisions based on current reality rather than outdated information.
MQTT Communication
Agents use an MQTT broker for communication. This decoupled system allows agents to talk to each other without knowing each other's specific addresses, making the coordination scalable and easy to add new modules.

Terminology

Summary

LLM-based agents are proposed as a solution for high-mix, low-volume manufacturing by generating production sequences and handling unforeseen runtime faults in flexible automation systems. The core contribution is the comparison of three agent architectures—orchestrator, peer-to-peer, and monolithic—using an MCP tool server exposing OPC UA skills to demonstrate a viable foundation for LLM-programmed smart manufacturing.

The gist: The monolithic and peer-to-peer architectures both achieve the highest mean solve rate (93%), while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment.

Agent Architecture and Tooling

The proposed solution pairs each factory module with a dedicated LLM-based agent and an MCP tool server that exposes the module’s skills via OPC UA method calls. The MCP (Model Context Protocol) is utilized because it introduces a standardized way to expose machine capabilities to LLMs, providing a structured interface through semantically enriched, self-describing tool interfaces rather than unstructured natural-language heuristics. The MCP server is the only component that knows the OPC UA endpoint address and method signatures, while the LLM agent knows only the tool’s name, description, and schema. This design enforces a separation of concerns central to skill-based manufacturing.

Multi-Agent Coordination via MQTT

Agents coordinate over an MQTT broker using publish/subscribe communication. This mechanism is chosen for its suitability in scalable multi-agent coordination, as agents are decoupled via a broker, allowing new agents to join without modifying existing ones, and topic-based routing supports one-to-many broadcast semantics. The three agent architectures compared are:

  1. Orchestrator architecture: A central agent holds a global view of the factory and delegates tasks to dedicated module agents.

  2. Peer-to-peer architecture: Module agents communicate directly with one another without a central coordinator, augmented with a send message tool for routing messages to peers.

  3. Monolithic architecture: A single agent controls all modules directly.

Real-Time State Injection and Grounding

The agents are grounded by real time updates of the factory state, which is externally injected into the system. This is achieved by tracking the effects of OPC UA method calls in an internal state tracker, which maintains a live snapshot of physical conditions like plate locations and belt occupancy. Each skill completion event received via MQTT updates this state. At each inference step, this state is passed to the LLM agents by prepending it to the input prompt as a structured context block, which relieves LLMs from having to reconstruct the state themselves from conversation history. Additionally, a Tool Availability Manager evaluates the current simulation state and publishes physically valid tools to each agent over MQTT, preventing impossible operations without requiring explicit reasoning about preconditions.

Experimental Results and Architectural Trade-offs

The system was evaluated across nine production challenges of increasing complexity, including silent hardware fault detection. The evaluation metrics include Solve Rate (SR), Time, Tokens consumed per run, and Tool Calls used.

"The monolithic and peer-to-peer architectures both achieved the highest mean solve rate (93%), while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment."

The results showed that for challenges 7, 8, and 9 (silent failures), the orchestrator performed worst on Challenge 7 (20% SR), where it entered a loop between agents without identifying the need to switch modules. Conversely, the monolithic architecture autonomously switched to an alternative module on all ten runs for those same fault scenarios. The ablation study confirmed that real-time factory state injection is critical for reliable fault diagnosis, as without it, the monolithic architecture’s solve rate drops from 93% to 88%.

Conclusion and Future Directions

The paper concludes that no single architecture dominates across all metrics. The monolithic architecture ties the peer-to-peer architecture for the highest mean SR of 93% and achieves the fastest execution time. The orchestrator excels at belt-routing challenges, achieving 100% on Challenge 9, but its centralized design is a single point of failure. Future work suggests investigating smaller, quantized models deployable on industrial edge devices and exploring shifting LLM agents from runtime controllers to production program generators rather than real-time tool callers against live hardware. Furthermore, the system's reliance on a local MQTT broker means it has not been evaluated under realistic network latency or bandwidth constraints. The hexagonal conveyor backbone also introduces physical centralization that constrains logical decentralization.

Index Terms

Multi-Agent System, Large Language Model, OPC UA, MQTT, Model Context Protocol, Industrial Automation, Smart Manufacturing, Cyber-Physical System.

The paper is the first work to integrate MCP tool servers with OPC UA in a multi-agent LLM framework for factory automation.

Improvements for AI systems

Here are the specific improvements and capabilities for an AI system derived from this research:


) Offline Production Sequence Generation: The LLM agent, deployed offline, can generate deterministic production sequences by reasoning over desired final product specifications and available modules/skills. This significantly reduces manual programming effort for complex, high-mix manufacturing tasks.

) Online Runtime Fault Detection and Autonomous Recovery: The system can operate live machines, continuously monitoring the factory state (via real-time MQTT injection). When a silent hardware fault occurs—where a machine reports success but produces no physical effect—the agent can reason over discrepancies between expected and actual states to attribute the failure, diagnose the root cause, and autonomously reroute production (e.g., switching modules or changing transport paths) without explicit pre-programmed failure logic.

) Standardized Modular Skill Integration: The system leverages the Model Context Protocol (MCP) to expose machine capabilities as semantically enriched, self-describing tool interfaces via OPC UA method calls. This allows the AI to interact with diverse manufacturing hardware using a standardized, structured interface rather than relying on unstructured natural language heuristics or proprietary APIs.

) Heterogeneous Agent Coordination: The system supports three distinct agent architectures—Orchestrator (centralized planning), Peer-to-Peer (decentralized coordination), and Monolithic (direct control)—allowing the AI to choose the most appropriate coordination strategy based on the required fault tolerance, reasoning scope, and communication overhead for a specific manufacturing challenge.

) Context-Aware Reasoning: Each LLM agent receives an up-to-date, structured context block at every inference step containing a real-time snapshot of the entire factory state (plate locations, belt occupancy, inventory). This enables agents to perform accurate reasoning and act locally while maintaining global awareness of the system's physical constraints.

) Dynamic Tool Constraint Management: The AI system incorporates a dynamic tool availability manager that evaluates the current simulation/physical state and publishes a filtered set of physically valid tools to each agent. This prevents the LLM from attempting operations that are physically impossible given the current state (e.g., attempting to load a plate into an unloaded module), ensuring operational safety and feasibility before execution.

) High-Fidelity Simulation for Training: The use of a Python-based simulation environment with mock OPC UA servers allows for training and testing agent behaviors across nine increasingly complex production challenges, including scenarios involving multi-type orders, block rotation, and silent faults, providing a robust testing ground for complex coordination strategies before physical deployment.

) Cost and Efficiency Optimization: The system provides metrics on token consumption per run broken down by agent. This allows researchers to optimize the LLM's prompting strategy—such as context-history trimming or task decomposition—to minimize operational costs while maximizing solve rate, providing a clear trade-off analysis between reasoning depth and computational efficiency.

Sources

Related papers