Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale".
Tom: Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re talking about "Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale," and the authors are Harsh Rao Dhanyamraju Leonidas Raghav, Aaron Lee, and others, who basically set out to look at how existing multi-agent systems fall short when you try to handle the massive scale of an enterprise.
Jane: The core idea is that current setups assume tasks come in one by one, but real enterprise AI deals with a constant stream of events—like alerts or user requests—so they needed a way for agents to react autonomously and continuously.
Lu: What’s striking about the title is that it explicitly mentions enterprise scale; it tells us this isn't just theoretical work on small prototypes, but something aimed at the massive environments where things actually happen every day.
Meng: When you look at those authors, they seem to have a solid background in both the engineering side and the data science side, which is important because tackling continuous streams requires both robust planning and reliable execution.
Lalam: I see a huge implication here for how we view AI systems; it’s suggesting that future AI shouldn't just be about answering questions but about being an always-on operational partner that manages its own workload intelligently.
The paper's summary: Tom: The authors summarize that they tested two main ways to run these agents: one called DAG Plan and Execute, which plans everything upfront, and another called ReAct, which just keeps reasoning and acting in a loop without a big initial plan.
Jane: They showed that the real challenge is agent discovery noise; when you have hundreds of agents, figuring out who to call next becomes the main headache for the system's performance.
Lu: That noise issue explains why they focused on comparing these two architectures across different sizes—Persona, Department, and Enterprise scales—to see which planning style handles that uncertainty better.
Meng: The summary also highlights a key addition: a Task Manager that acts like an operating system scheduler for the AI, deciding when to start a new job or combine incoming events into one task.
Lalam: This Task Manager sounds incredibly powerful because it handles the chaos of simultaneous events by prioritizing them and even merging related pieces of information together so the system doesn't get bogged down in duplicate work.
The paper's improvements: Tom: The authors suggest that this orchestration system improves things by adding that Task Manager, which lets the system decide if an incoming event should be a new task or merged into an existing one based on priority and context.
Jane: They also found that using this manager to handle preemption—interrupting a current task for a more critical one—and merging related events significantly boosts the correctness of those long-running operations.
Lu: They also introduced improvements in how the DAG Plan and Execute approach works, suggesting it can be very precise with parallel work when you’re dealing with smaller agent sets and well-defined dependencies.
Meng: For practical deployment, they emphasize that this whole system needs deterministic guardrails to make sure the AI's decisions about merging or preemption don't accidentally violate actual business rules.
Lalam: I think the ability to perform related-event merging is huge for culture because it means the system can actually understand the bigger picture of an incident, not just process individual alerts in isolation.
Conclusion: Tom: To wrap things up, this paper shows that scale is what really matters for orchestration performance because agent discovery noise gets worse as you go from small teams to huge enterprises.
Jane: They concluded that the Task Manager provides a real boost, showing significant improvements in handling high-priority queues and making related events more coherent at enterprise levels.
Lu: The paper suggests that the choice between DAG Plan and Execute and ReAct should really depend on whether your agent set is small enough for structured planning or if you need the robustness of ReAct when discovery is noisy at scale.
Meng: I think the practical implication for us right now is that we have to focus on monitoring signals like backlog age and preemption frequency to know exactly where our system's bottleneck is.
Lalam: I just feel really optimistic about this direction because having an AI that can manage continuous streams with priority and merging capabilities means we are closer to having truly proactive digital workers.
SAP SE
cs.AI, cs.MA
Submitted: 2026-06-18
Updated: 2026-09-28
Code: https://github.com/langchain-ai/langgraph
Project page: https://google.github.io/adk-docs
Importance score: 86/100
The gist: Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and
Key concepts
- DAG Plan & Execute
- This architecture requires a planner to create a complete execution graph before any action starts. It separates planning from execution, which can be precise for structured workflows but suffers from overhead when replanning is needed at large enterprise scales.
- ReAct
- ReAct uses a single model that continuously interleaves reasoning and taking actions in a loop without upfront planning. This approach is more robust when agent discovery is noisy because it avoids committing to a single, potentially flawed global plan.
- Task Manager
- This component acts like an operating system scheduler for events. It decides whether to start new tasks or merge incoming events based on context and priority, managing tasks through states like QUEUED or ACTIVE to enable continuous operation.
Terminology
Summary
Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale. The gist: scale, not task complexity, dominates orchestration performance as agent discovery noise becomes the primary bottleneck in enterprise settings.
Orchestration Architectures Compared
The paper evaluates two primary orchestration architectures across 208 production-derived scenarios spanning Persona (<10 agents), Department (20–80 agents), and Enterprise (200) scales: DAG Plan & Execute and ReAct. The core difference lies in their approach to planning: DAG Plan & Execute uses a structured approach where a Planner generates an execution graph upfront,
separating planning from execution, while ReAct is the default architecture where a single model interleaves reasoning and action in a continuous loop without upfront planning.
The Role of the Task Manager
A critical component introduced to support continuous operation is the Task Manager, which functions analogously to a process scheduler in an operating system kernel.
This manager handles incoming events by leveraging LLM-powered decision making to determine whether to create a new task or merge it into an existing one, and assigns the appropriate priority level based on event content and context.
It manages tasks through states like QUEUED, ACTIVE, and COMPLETED. The Task Manager integrates with both architectures at defined stoppage points—checkpoints in DAG Plan & Execute or after an agent call returns in ReAct—to enable preemption or task updates (from merged events).
Benchmark Design and Scenario Taxonomy
The evaluation dataset comprises 208 scenarios derived from actual user stories, structured into five types to test specific orchestration capabilities:
-
Simple scenarios (1–3 agents) establish baseline competency.
-
Parallel scenarios (4–8 agents) test concurrent execution of independent subproblems.
-
Complex scenarios (5–10 agents) combine dependency handling with parallel work, creating multi-stage execution paths where upstream agents provide entity identifiers required by downstream agents while independent branches can proceed concurrently.
-
Failure scenarios introduce incomplete or ambiguous responses, requiring the orchestrator to
detect missing information, retry with refined requests, or proceed with partial data.
-
Priority and related-event scenarios test the Task Manager’s continuous operation capabilities, verifying
correct triage when critical events (equipment failures, safety incidents) arrive during lower-priority work
and requiring systems tomerge related events into a unified task.
Performance Findings Across Scales
The results show that scale, not task complexity, dominates orchestration performance,
as both architectures degrade at enterprise scale due to agent discovery noise. ReAct is noted as the stronger default when discovery is noisy,
while DAG Plan & Execute shows advantages on structured workflows at small-to-medium scales through better parallelization and precision. The Task Manager demonstrated significant value, reducing high-priority queue latency by 14–75%
and improving related-event correctness by over 20 percentage points at enterprise scale.
Enterprise Deployment Implications
The choice between architectures should be driven less by task complexity than by registry scale and recovery cost. DAG Plan & Execute is attractive when the relevant agent set is small or moderately sized, but its plan-interrupt-replan overhead compounds errors at enterprise scale.
ReAct is the stronger default when discovery is noisy, as its stepwise interaction gives it fewer opportunities to overcommit to an incorrect global plan,
making it more robust at 200-agent scale. The Task Manager's value is most pronounced when the workload is continuous, enabling a scheduler for event streams through priority inference and preemption. Furthermore, the paper emphasizes that safety decisions must remain with workflow-specific specialist agents rather than the orchestrator itself.
Key Metrics and Analysis
The evaluation used LLM-as-a-Judge evaluation to score correctness by comparing the orchestrator’s final synthesized response against expected answer fragments.
Precision and recall analysis revealed that DAG achieves higher precision at every scale as upfront planning selects agents more carefully, but ReAct’s recall holds up better because it lacks the machinery that compounds errors.
Cost analysis showed that all systems converge toward higher consumption (20–55K per event) at Enterprise scale as discovery retries and replanning dominate. The findings confirm that the bottleneck at scale is agent selection rather than the orchestration architecture itself.
Limitations of the Study
The study notes several limitations, including mock agent latency, which compresses queue dynamics,
fixed orchestration parameters, and the exclusion of persistent memory to establish a controlled baseline independent of learned optimizations. Additionally, all experiments assume a single orchestration worker; future work is needed to address distributed scheduling challenges. The paper concludes that enterprise systems should monitor signals like backlog age, preemption frequency, merge accuracy, repeated agent calls, and discovery recall
to diagnose the primary bottleneck.
Improvements for AI systems
Here are specific improvements for AI systems based on the scientific paper, detailing what the improved system can achieve:
-
The introduction of a dedicated, LLM-powered
Task Manager
layer fundamentally shifts enterprise AI from reactive chatbots to autonomous digital workers capable of continuous operation. -
This Task Manager enables high-priority event preemption and related-event merging, allowing the system to handle asynchronous event streams (alerts, user requests) rather than discrete request-response cycles.
-
The improved system can maintain a deterministic priority queue (LOW, MEDIUM, HIGH) that prevents task starvation by incorporating task age and progress factors into its scoring function.
-
The AI can perform
related-event merging,
synthesizing fragmented events pertaining to the same underlying situation into a single unified task, significantly improving context coherence for long-horizon operations. -
For enterprise scales (200 agents), the system utilizes a flexible orchestration choice:
-
If the agent discovery space is noisy (high scale), it defaults to the ReAct architecture, which handles failures incrementally and avoids costly full replanning cycles when faced with unreliable agent returns.
-
If tasks involve structured workflows with known dependencies at small-to-medium scales, it utilizes the DAG Plan & Execute architecture, leveraging upfront planning to ensure precision and structured parallelization.
-
The system can dynamically adjust execution strategy based on scale: ReAct is preferred for robustness against discovery noise at enterprise scale, while DAG excels where explicit parallelization and dependency mapping are beneficial at smaller scales.
-
The improved system incorporates
two-tier agent discovery
(lightweight summaries first, full cards on demand), ensuring the planning phase remains computationally bounded even with hundreds of registered agents. -
The AI can perform
deterministic guardrails
by validating LLM decisions against configuration constraints before mutating the Task Backlog, ensuring no event is silently dropped and actions remain within operational/business boundaries. -
The system separates orchestration from safety: the orchestrator schedules work and routes context, while specialized domain agents (e.g., a Finance agent) retain ultimate authority for enforcing workflow-specific business rules (e.g., segregation of duties checks).
-
The AI can generate precise execution plans using the DAG Planner's two-phase approach: first, determining necessary information via a directed reasoning step, and second, constructing the execution graph by strictly enforcing dependency minimization and requiring explicit entity identifiers in node descriptions.
-
The system can dynamically adapt its plan through
checkpoint
mechanisms—pausing execution before replanning to inject merged event context or preempt the current task for higher-priority work—ensuring continuous responsiveness. -
The AI can use LLM-as-a-Judge evaluation, comparing the orchestrator’s final synthesized response against expected answer fragments derived from agent guidelines, providing a quantifiable correctness score based on semantic presence of required information.
-
The improved system can translate complex process mining data into prescriptive actions using the Dashboard Analyzer Agent, allowing it to autonomously interpret event logs and recommend next best actions for continuous process optimization.
Sources
- The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption
- DynTaskMAS: A Dynamic Task Graph-driven Framework for Asynchronous and Parallel LLM-based Multi-Agent Systems
- Multi-Agent Collaboration via Evolving Orchestration
- Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents
- From Agent-Only Social Networks to Autonomous Scientific Research: Lessons from OpenClaw and Moltbook, and the Architecture of ClawdLab and Beach.Science
- Autonomous Deep Agent
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection