Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing

arXiv:2608.12371 · cs.AI · Submitted 2026-07-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing".

Jane: The paper was written by Sabeur Lajili and Zaki Brahmi from University of Sousse.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1 — Tom and Jane discuss title and authors of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Tom: Jane, we've got a real treat today. The paper is called "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and it's from Sabeur Lajili and Zaki Brahmi at the University of Sousse in Tunisia.

Jane: And I have to say, Tom, that title is a mouthful, but it's actually describing something we all rely on every day. Think about your smartwatch tracking your heart rate, or a city's traffic cameras analyzing intersections in real time. All of that is stream processing, and it needs to happen fast.

Tom: Right, and the problem is where do you run those computations? You can't send everything to a giant cloud data center because the latency would kill you. So you push it to the edge, to small clusters of computers physically close to the sensors.

Jane: Exactly. But now you have dozens of these edge clusters, and they're all making decisions independently. It's like a bunch of restaurants in a city all guessing how many customers will show up tonight, without talking to each other. Some will over-order, some will under-order.

Tom: And that's where the "Contract Net Protocol" comes in. It's an old idea from the 1980s, basically a formal way for agents to announce a task, receive bids, and award the work. But the authors here, they've supercharged it with large language models.

Jane: So instead of just saying "I need two CPUs and four gigs of RAM," the agent can say something like, "I need to migrate this high-priority ECG stream, and my reliability monitor just flagged a traffic spike near you." The LLM helps interpret that messy, real-world context.

Tom: The authors call it LLM-MR-CNP. And the key insight is that they keep the hard math, the resource checks, the deadlines, all deterministic. The LLM is only used for the fuzzy stuff, like interpreting warnings and refining proposals.

Jane: I love that separation. It's like having a brilliant negotiator who can read the room and adjust their pitch, but they still have to get the contract signed by a lawyer who checks every number. The LLM suggests, the rules decide.

Tom: And that's the part that gets me excited. We've seen so many papers where people just let the LLM loose and hope for the best. This one is careful. It's a hybrid, and that's what makes it practical.

Jane: Practical is the word. Because if you're running a real edge network, you can't have an AI hallucinate a resource allocation and crash a hospital's monitoring system. You need those guardrails.

Tom: Absolutely. And the authors tested this against real Alibaba cluster traces, which we'll get into in a minute. But first, I want to say, this is the kind of paper that bridges the gap between the AI hype and actual systems engineering.

Jane: It really does. And I think the biggest implication is that we don't have to choose between smart AI and reliable systems. We can have both, if we design the boundaries carefully.

Tom: Stay with us, because next we're going to break down exactly what they did in the experiments and why the results are so promising.

Paper discussion segment 2 — Tom and Jane discuss the paper's summary of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Jane: Welcome back. We're still on "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and Tom, I want to dig into what they actually found.

Tom: The headline numbers are striking. In their first experiment, they simulated a workload spike, a data drift scenario. The classic single-round approach had a latency violation rate of zero point five three. That means over half the time, tasks missed their deadlines.

Jane: That's terrible. And then they added multi-round negotiation, just the rule-based version, no LLM. That dropped the violation rate to zero point three seven. So just letting agents talk back and forth, revise their bids, that alone helped a lot.

Tom: Right. But then they added the LLM-assisted layer, and the violation rate collapsed to zero point zero three. Three percent. That's a massive improvement, from broken to nearly perfect.

Jane: And the utility score went from one point zero three to one point six one. So not only are they meeting deadlines, they're making better overall decisions about where to put the work.

Tom: Now, Jane, here's the part that really matters for real systems. They also tested concurrent requests, like twenty agents all trying to place twelve tasks at the same time. The single-round approach overcommitted resources. It promised more CPU and memory than actually existed.

Jane: That's the classic tragedy of the commons. Everyone sees the same free resources, everyone grabs them, and then everyone crashes.

Tom: Exactly. But both multi-round approaches, with and without the LLM, eliminated overcommitment entirely. Zero. The negotiation process itself, reserving capacity as you go, that fixed the problem.

Jane: So the multi-round structure is doing the heavy lifting, and the LLM is adding a smaller but real boost. They saw conflict resolution go from zero point eight six to zero point nine one with twenty agents, and utility improved by up to twenty-two percent.

Tom: And that's the honest finding. The LLM isn't magic. It's a refinement tool. The protocol change, going from one round to multiple rounds, that's the foundation.

Jane: It reminds me of how human teams work. Just having a meeting where people can change their minds after hearing new information is huge. The LLM is like having a really good facilitator in that meeting.

Tom: But there's a cost. The LLM version took eighty-three seconds of reasoning time in the drift scenario, and it used three times as many messages. For a latency-sensitive system, that's real overhead.

Jane: So you wouldn't want to invoke the LLM every single time. You'd want to use it selectively, only when the situation is ambiguous or high-stakes.

Tom: That's exactly what the authors suggest. Routine decisions stay deterministic, and you escalate to the LLM when there's a reliability warning or conflicting proposals. It's an adaptive strategy.

Jane: And I think that's the most practical takeaway. You don't need AI everywhere. You need AI where the rules can't capture the nuance.

Tom: Next up, we're going to look at the third experiment, where they compared different LLMs and prompting strategies. That's where things get really interesting, because not all models are created equal.

Paper discussion segment 3 — Tom and Jane discuss the improvements the paper suggests of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Jane: So Tom, in this third experiment, they took twenty-five test cases and threw a bunch of different LLMs at the problem. This is where the paper gets really practical about model choice.

Tom: And the results are fascinating. They compared a small local LLM against a larger one. The small one got a CFP accuracy of zero point four six, which means it struggled to even formulate the call for proposals correctly. The larger model got zero point five six.

Jane: But here's the twist. The larger model, when used in a single-round negotiation, only reached zero point seven one offloading accuracy. That means it picked the right destination less often than the small model did in multi-round mode, which hit zero point seven eight.

Tom: So model size alone doesn't win. The interaction structure matters more. When they let the larger model negotiate over multiple rounds, it jumped to zero point eight eight. That's a huge leap from zero point seven one.

Jane: It's like the difference between asking one person for directions and asking three people, then comparing their answers. The back-and-forth is what catches mistakes.

Tom: Exactly. And then they went further and tested five different LLMs with six different prompting strategies. Zero-shot, few-shot, chain-of-thought, ReAct, and combinations. Thirty configurations total.

Jane: And the winner? Gemini-three-Flash with a ReAct plus few-shot plus chain-of-thought hybrid hit a perfect one point zero zero offloading accuracy. But it wasn't the fastest or cheapest.

Tom: Right. And here's the warning in the data. For DeepSeek-V4-Pro, adding more prompting components actually hurt. It went from zero point seven two accuracy down to zero point six eight. More complexity isn't always better.

Jane: That's such an important finding. You can't just stack prompting techniques like Lego bricks. Each model has its own personality, and you have to find what works for that specific model.

Tom: And they also measured tokens and latency. The small local Llama3 model took one hundred eighty-six seconds per case. That's way too slow for real-time scheduling. The larger models were faster, around twenty to seventy-five seconds.

Jane: So there's a real trade-off between accuracy, speed, and cost. A perfect one point zero zero accuracy is great, but if it takes too long, the stream might miss its deadline anyway.

Tom: The authors also flagged that these results are descriptive, not definitive. Twenty-five cases is a small sample. One changed prediction shifts accuracy by four percentage points.

Jane: Still, the pattern is clear. You need to match the model and the prompt to the specific negotiation task. There's no one-size-fits-all.

Tom: And I think the biggest improvement this paper suggests is the idea of selective escalation. Don't run the LLM on every decision. Run it only when the deterministic rules hit a conflict or an ambiguity.

Jane: That's the smart engineering approach. Use the expensive, powerful tool only when you actually need it. That's how you get the best of both worlds.

Tom: And that's exactly where we're heading in the conclusion, because this has big implications for how we build real systems.

Conclusion — Tom and Jane summarize the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications, say goodbye to the paper and get ready to discuss the next one.: Jane: Alright, Tom, let's wrap this up. We've been talking about "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and I think the core message is really elegant.

Tom: It is. The authors took a forty-year-old protocol, the Contract Net Protocol, and showed that the biggest gains come from just letting agents negotiate over multiple rounds. That alone cut latency violations in half and eliminated resource overcommitment.

Jane: And then the LLM adds the cherry on top. It handles the messy, qualitative context, the reliability warnings, the uncertain forecasts, and it pushes performance even higher.

Tom: But they were honest about the costs. The LLM is slower, it uses more tokens, and it's not always better. The best model-prompt combination hit perfect accuracy, but another model got worse when they added more prompting tricks.

Jane: So the real takeaway is about boundaries. Keep the hard constraints deterministic. Let the LLM handle interpretation and refinement. And escalate to the LLM selectively, only when the situation genuinely needs it.

Tom: And that's a blueprint we can actually build on. This isn't a paper that just says "AI is amazing." It says "here's exactly where AI helps, and here's where it doesn't."

Jane: I also appreciate that they released their benchmark, prompts, and raw outputs. That means other researchers can replicate and build on this work. That's how science moves forward.

Tom: For the real world, think about smart cities, autonomous vehicles, industrial IoT. Anywhere you have latency-sensitive streams running on distributed edge resources. This approach could make those systems more reliable without requiring a complete redesign.

Jane: And I think that's the most exciting part. You don't have to throw away your existing scheduling system. You just extend it with negotiation rounds and selectively add an LLM where it helps.

Tom: Well said, Jane. So we're saying goodbye to this paper, but we're taking its lessons with us. Next up, we've got a paper on federated learning for medical imaging that I think is going to spark some debate.

Jane: Can't wait. Thanks for listening, everyone. We'll see you on the next episode.

Sabeur Lajili, Zaki Brahmi

University of Sousse

cs.AI

Submitted: 2026-07-19

Comments: 8

Code: https://github.com/MythesisProject2024/MAS

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 61/100

Key concepts

Stream Processing
Processing data continuously as it is generated, such as tracking heart rates or analyzing traffic cameras in real time. This type of processing requires fast execution to avoid critical delays (latency).
Mobile Edge Computing
A system where computations are pushed to small clusters of computers physically close to the sensors (the 'edge') rather than relying on distant cloud data centers, minimizing latency.
Contract Net Protocol
An established protocol for multiple agents to negotiate tasks. Agents announce a task, receive bids from others, and award the work. The authors supercharge this with LLMs.
LLM-Assisted Negotiation
Using Large Language Models (LLMs) not for core math or resource checks, but for interpreting messy, real-world context (like reliability warnings) and refining proposals during the negotiation process.

Terminology

Summary

Summary

This paper introduces MAS-DecStream, a multi-agent scheduling system for stream processing in mobile edge computing, whose main contribution is LLM-MR-CNP: an extension of the classical Contract Net Protocol (CNP) with semantic CFP formulation, progressive context disclosure, multi-round proposal revision, negotiation memory, and deterministic validation. The system addresses the challenge that independent local decisions can therefore select the same destination, overload scarce resources, delay high-priority streams, and violate QoS constraints in heterogeneous edge–cloud infrastructures.

The proposed architecture consists of two interacting layers: an execution layer that monitors runtime telemetry, predicts near-future load, and applies migration decisions, and a collaborative layer containing stateful edge-cluster agents orchestrated using LangGraph. Each agent combines local observation and predicted state, task and QoS knowledge, task-scoped negotiation memory, an LLM for contextual interpretation and message generation, and tools for state retrieval, requirement validation, and utility computation. Critically, the LLM may interpret warnings and refine a proposal, but measurements, hard constraints, and numerical ranking remain deterministic.

LLM-MR-CNP extends classical CNP through five steps: (1) an overloaded initiator broadcasts a compact CFP containing essential task demand and scheduling intent; (2) each responder inspects its current and predicted state and returns a proposal or refusal; (3) the initiator discards responses that violate hard constraints; (4) when several feasible candidates remain close or a proposal is uncertain, the initiator progressively discloses additional context such as forecast load, task priority, or reliability warnings, and requests revised proposals only from remaining candidates; (5) negotiation terminates when one feasible candidate remains, a stable best candidate emerges, or the maximum number of rounds is reached. The final allocation is determined by the equation A* = A j in A feasible(T k) U j(T k), where the selected destination maximizes utility among feasible candidates.

The scheduling objective is formalized as (D*, r*) = D,r OF(D, M r), where OF(D, M r) = alpha 1 Latency(D) + alpha 2 Energy(D) + alpha 3 (1 - LBD total(D)) + alpha 4 CO(M r), with weights satisfying alpha m at least 0 and sum m alpha m = 1. The term LBD total(D) in [0,1] denotes the overall load-balancing degree, and CO(M r) denotes coordination overhead including the number and size of exchanged messages.

The experimental methodology compares three scheduling pipelines: RB-SR-CNP (classical one-round announcement–proposal–award exchange), RB-MR-CNP (iterative proposal refinement with updated quantitative forecasts but no LLM), and MAS-DecStream (multi-round workflow with LLM-assisted CFP interpretation, qualitative-context handling, and natural-language proposal revision). All conditions receive the same task instances, candidate clusters, current resource states, hard feasibility rules, and deterministic utility function. The workload is derived from the Alibaba ASI Trace 2026 job-execution summary, sampling 1,000 records enriched with stream- and MEC-specific attributes, yielding 5,000 task–cluster combinations across five heterogeneous cluster profiles (high-compute, energy-efficient, low-latency, privacy-enabled, and overloaded-source).

The paper addresses three research questions with the following findings:

RQ1 (Migration under data drift): An unmodelled workload spike creates overload and deadline risk over a 300-second horizon sampled every 10 seconds, with migration triggered in 23 of 30 windows. RB-SR-CNP has the lowest coordination cost but produces latency violations in 0.53 of windows. Replacing this with rule-based refinement (RB-MR-CNP) reduces the violation rate to 0.37, confirming iterative proposal revision is useful even without an LLM. MAS-DecStream further reduces the rate to 0.03 and achieves the highest utility (1.61 versus 1.11 and 1.03), but at the cost of 276 messages (versus 92 and 184) and 83.56 seconds of LLM reasoning time. The authors note this motivates selective invocation rather than using an LLM at every monitoring window.

RQ2 (Concurrent conflict resolution): The setting scales from 5 agents with 3 requests to 20 agents with 12 concurrent requests, including ECG monitoring, video surveillance, and emergency alarms. RB-SR-CNP makes awards without sufficient reconciliation, leading to persistent overcommitment and negative utility. Both multi-round methods eliminate overcommitment in every setting, demonstrating that successive validation, reservation, and proposal refinement constitute the principal source of robustness under concurrency. MAS-DecStream adds smaller but consistent gains: utility improves by approximately 22%, 12%, and 4% for the 5-, 10-, and 20-agent settings respectively, while conflict resolution increases by five percentage points at 10 and 20 agents (from 0.90 to 0.95 and 0.86 to 0.91). Decision time increases from 700 to 840 ms at 10 agents and from 1600 to 1950 ms at 20 agents, leading the authors to conclude LLM assistance is most defensible for ambiguous, high-priority, or context-dependent requests.

RQ3 (LLM and prompting effects): Over 25 cases, a small local LLM achieves 0.46 CFP accuracy and 0.78 offloading accuracy, while a larger LLM under single-round negotiation reaches only 0.71 offloading accuracy despite 0.56 CFP accuracy. Allowing multi-round refinement raises the large model's offloading accuracy to 0.88, indicating the interaction structure is at least as important as raw model capability. This improvement comes with substantial cost: token consumption grows from 19,100 to 53,452 and decision time rises from 20.10 to 75.15 seconds. The large multi-round model is six points above RB-MR-CNP in final-host accuracy (0.88 versus 0.82), while responder accuracy remains nearly unchanged, suggesting the main benefit is not uniformly better local accept/refuse judgments, but the initiator's ability to combine, refine, and validate multiple imperfect responses before the final award.

The prompting study evaluates zero-shot, few-shot, CoT, ReAct, and hybrid prompts across five LLMs (GPT-OSS:20B, DeepSeek-V4-Pro, GLM-5.2, Gemini-3-Flash, and Llama3). Key findings include: the best prompting strategy is model-dependent; few-shot and hybrid prompts frequently improve CFP coverage but a better CFP does not necessarily produce a better final destination; Gemini-3-Flash reaches the highest observed offloading accuracy of 1.00 with ReAct+few-shot+CoT; CoT+Few-shot provides the best balance of accuracy and token cost; and ReAct-based prompting incurs higher costs without consistently improving performance. The authors caution that each configuration contains only 25 cases and one run, the values are treated as descriptive observations rather than stable estimates of general model superiority.

The paper identifies three main findings: (1) multi-round negotiation consistently provides the largest performance gain, reducing drift violations and eliminating overcommitment even without LLM assistance; (2) the agentic-AI layer is most beneficial when decisions involve qualitative, ambiguous, or partially structured information; (3) deterministic verification remains essential, as Agentic AI generates and refines negotiation proposals, whereas final scheduling decisions are validated against resource, QoS, and utility constraints before execution.

The authors acknowledge threats to validity: construct validity concerns about CFP accuracy measurement; internal validity limitations because Scenarios 1 and 2 compare complete pipelines and their difference is not a causal estimate of the LLM component; conclusion validity issues due to 25 cases and one run per configuration; and external validity limitations because workloads are enriched from a production AI trace rather than executed in a physical deployment, with the largest configuration including only 20 agents and 12 concurrent requests.

The paper concludes that extending the single proposal–award cycle to multiple validated rounds provides the largest and most consistent gain, while LLM assistance adds smaller but useful improvements when refinement depends on ambiguous or qualitative runtime context and exhibits model-dependent accuracy–cost trade-offs. The study is positioned as an initial empirical assessment of a hybrid agentic scheduling architecture rather than proof of general superiority over established schedulers. Future work will evaluate selective LLM escalation, larger decentralized deployments, and stateful stream-operator migration under measured network and recovery costs.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:

Implementation: Replace pure LLM reasoning with a two-layer system where an LLM handles qualitative interpretation and natural-language negotiation, while deterministic tools enforce hard constraints (resource capacity, deadlines, feasibility).

What it can do:

  • Eliminate resource overcommitment entirely (0% in all tested configurations, vs. 5–20% for single-round approaches)

  • Reduce latency violations from 53% to 3% under workload drift

  • Guarantee that LLM-generated proposals are always validated against measurable system state before execution

Implementation: Extend single-shot decision-making with bounded iterative negotiation: initial proposal → progressive context disclosure → targeted revision requests → deterministic final validation.

Implementation: Route routine scheduling decisions through deterministic rules; activate LLM-assisted negotiation only when: (a) proposals are uncertain/conflicting, (b) qualitative context (reliability warnings, privacy constraints) is present, or (c) high-priority tasks are at risk.

Implementation: Instead of accumulating prompt components indiscriminately, select prompting strategy per model based on measured accuracy–cost trade-offs.

Implementation: Initially broadcast minimal task information; only disclose additional context (forecast load, priority, reliability warnings) to unresolved candidates in later rounds.

Implementation: Maintain per-task interaction history across rounds, enabling agents to track provisional proposals, revisions, and withdrawals without cross-task contamination.

Implementation: Terminate negotiation when: (a) one feasible candidate remains, (b) the leading candidate is sufficiently separated, (c) the decision stabilizes, or (d) maximum rounds reached. Reject or replace malformed LLM outputs with deterministic fallbacks.

The improved AI system can:

  • Schedule latency-sensitive stream tasks across heterogeneous edge–cloud infrastructures with 97% deadline compliance

  • Resolve concurrent resource conflicts with 91% success at scale (20 agents)

  • Adapt to workload drift without centralized coordination

  • Explain its decisions through natural-language negotiation traces

  • Operate cost-effectively by escalating to LLM reasoning only when context demands it

  • Guarantee hard constraints through deterministic validation layers

This system is particularly suited for smart-city infrastructure, industrial IoT, and real-time analytics where both qualitative judgment and hard resource guarantees are required.

Abstract

Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes MAS-DecStream, whose main contribution is LLM-MR-CNP: an extension of the classical Contract Net Protocol with semantic CFP formulation, progressive context disclosure, multi-round proposal revision, negotiation memory, and deterministic validation. Edge-cluster agents refine natural-language offloading proposals from local observations, predicted resource states, and qualitative runtime context, while hard resource and QoS constraints remain deterministic. Experiments derived from the Alibaba ASI Trace evaluate the extension at three levels: single- versus multi-round CNP, rule-based versus LLM-assisted refinement, and fixed-model single- versus multi-round negotiation. Under the evaluated configurations, MAS-DecStream reduces latency violations to 3%, eliminates resource overcommitment, reaches a conflict-resolution rate of 0.91 with 20 agents, and improves utility by up to 22% over the multi-round rule-based baseline. A separate 25-case evaluation shows model- and prompt-dependent accuracy--cost trade-offs. The results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.

Sources

Related papers