Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing

summary

Video file (mp4)

In short

The episode discusses 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing.' Hosts analyze how combining a multi-round negotiation protocol with selective Large Language Model (LLM) assistance improves resource allocation and reduces latency violations in complex, real-world edge computing systems.

Key concepts

Stream Processing
Processing data continuously as it is generated, such as tracking heart rates or analyzing traffic cameras in real time. This type of processing requires fast execution to avoid critical delays (latency).
Mobile Edge Computing
A system where computations are pushed to small clusters of computers physically close to the sensors (the 'edge') rather than relying on distant cloud data centers, minimizing latency.
Contract Net Protocol
An established protocol for multiple agents to negotiate tasks. Agents announce a task, receive bids from others, and award the work. The authors supercharge this with LLMs.
LLM-Assisted Negotiation
Using Large Language Models (LLMs) not for core math or resource checks, but for interpreting messy, real-world context (like reliability warnings) and refining proposals during the negotiation process.

Terminology used across episodes

This episode discusses

The paper

Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing · Read on arXiv

Sabeur Lajili, Zaki Brahmi

University of Sousse

Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes MAS-DecStream, whose main contribution is LLM-MR-CNP: an extension of the classical Contract Net Protocol with semantic CFP formulation, progressive context disclosure, multi-round proposal revision, negotiation memory, and deterministic validation. Edge-cluster agents refine natural-language offloading proposals from local observations, predicted resource states, and qualitative runtime context, while hard resource and QoS constraints remain deterministic. Experiments derived from the Alibaba ASI Trace evaluate the extension at three levels: single- versus multi-round CNP, rule-based versus LLM-assisted refinement, and fixed-model single- versus multi-round negotiation. Under the evaluated configurations, MAS-DecStream reduces latency violations to 3%, eliminates resource overcommitment, reaches a conflict-resolution rate of 0.91 with 20 agents, and improves utility by up to 22% over the multi-round rule-based baseline. A separate 25-case evaluation shows model- and prompt-dependent accuracy--cost trade-offs. The results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing".

Jane: The paper was written by Sabeur Lajili and Zaki Brahmi from University of Sousse.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1 — Tom and Jane discuss title and authors of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Tom: Jane, we've got a real treat today. The paper is called "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and it's from Sabeur Lajili and Zaki Brahmi at the University of Sousse in Tunisia.

Jane: And I have to say, Tom, that title is a mouthful, but it's actually describing something we all rely on every day. Think about your smartwatch tracking your heart rate, or a city's traffic cameras analyzing intersections in real time. All of that is stream processing, and it needs to happen fast.

Tom: Right, and the problem is where do you run those computations? You can't send everything to a giant cloud data center because the latency would kill you. So you push it to the edge, to small clusters of computers physically close to the sensors.

Jane: Exactly. But now you have dozens of these edge clusters, and they're all making decisions independently. It's like a bunch of restaurants in a city all guessing how many customers will show up tonight, without talking to each other. Some will over-order, some will under-order.

Tom: And that's where the "Contract Net Protocol" comes in. It's an old idea from the 1980s, basically a formal way for agents to announce a task, receive bids, and award the work. But the authors here, they've supercharged it with large language models.

Jane: So instead of just saying "I need two CPUs and four gigs of RAM," the agent can say something like, "I need to migrate this high-priority ECG stream, and my reliability monitor just flagged a traffic spike near you." The LLM helps interpret that messy, real-world context.

Tom: The authors call it LLM-MR-CNP. And the key insight is that they keep the hard math, the resource checks, the deadlines, all deterministic. The LLM is only used for the fuzzy stuff, like interpreting warnings and refining proposals.

Jane: I love that separation. It's like having a brilliant negotiator who can read the room and adjust their pitch, but they still have to get the contract signed by a lawyer who checks every number. The LLM suggests, the rules decide.

Tom: And that's the part that gets me excited. We've seen so many papers where people just let the LLM loose and hope for the best. This one is careful. It's a hybrid, and that's what makes it practical.

Jane: Practical is the word. Because if you're running a real edge network, you can't have an AI hallucinate a resource allocation and crash a hospital's monitoring system. You need those guardrails.

Tom: Absolutely. And the authors tested this against real Alibaba cluster traces, which we'll get into in a minute. But first, I want to say, this is the kind of paper that bridges the gap between the AI hype and actual systems engineering.

Jane: It really does. And I think the biggest implication is that we don't have to choose between smart AI and reliable systems. We can have both, if we design the boundaries carefully.

Tom: Stay with us, because next we're going to break down exactly what they did in the experiments and why the results are so promising.

Paper discussion segment 2 — Tom and Jane discuss the paper's summary of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Jane: Welcome back. We're still on "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and Tom, I want to dig into what they actually found.

Tom: The headline numbers are striking. In their first experiment, they simulated a workload spike, a data drift scenario. The classic single-round approach had a latency violation rate of zero point five three. That means over half the time, tasks missed their deadlines.

Jane: That's terrible. And then they added multi-round negotiation, just the rule-based version, no LLM. That dropped the violation rate to zero point three seven. So just letting agents talk back and forth, revise their bids, that alone helped a lot.

Tom: Right. But then they added the LLM-assisted layer, and the violation rate collapsed to zero point zero three. Three percent. That's a massive improvement, from broken to nearly perfect.

Jane: And the utility score went from one point zero three to one point six one. So not only are they meeting deadlines, they're making better overall decisions about where to put the work.

Tom: Now, Jane, here's the part that really matters for real systems. They also tested concurrent requests, like twenty agents all trying to place twelve tasks at the same time. The single-round approach overcommitted resources. It promised more CPU and memory than actually existed.

Jane: That's the classic tragedy of the commons. Everyone sees the same free resources, everyone grabs them, and then everyone crashes.

Tom: Exactly. But both multi-round approaches, with and without the LLM, eliminated overcommitment entirely. Zero. The negotiation process itself, reserving capacity as you go, that fixed the problem.

Jane: So the multi-round structure is doing the heavy lifting, and the LLM is adding a smaller but real boost. They saw conflict resolution go from zero point eight six to zero point nine one with twenty agents, and utility improved by up to twenty-two percent.

Tom: And that's the honest finding. The LLM isn't magic. It's a refinement tool. The protocol change, going from one round to multiple rounds, that's the foundation.

Jane: It reminds me of how human teams work. Just having a meeting where people can change their minds after hearing new information is huge. The LLM is like having a really good facilitator in that meeting.

Tom: But there's a cost. The LLM version took eighty-three seconds of reasoning time in the drift scenario, and it used three times as many messages. For a latency-sensitive system, that's real overhead.

Jane: So you wouldn't want to invoke the LLM every single time. You'd want to use it selectively, only when the situation is ambiguous or high-stakes.

Tom: That's exactly what the authors suggest. Routine decisions stay deterministic, and you escalate to the LLM when there's a reliability warning or conflicting proposals. It's an adaptive strategy.

Jane: And I think that's the most practical takeaway. You don't need AI everywhere. You need AI where the rules can't capture the nuance.

Tom: Next up, we're going to look at the third experiment, where they compared different LLMs and prompting strategies. That's where things get really interesting, because not all models are created equal.

Paper discussion segment 3 — Tom and Jane discuss the improvements the paper suggests of the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications.: Jane: So Tom, in this third experiment, they took twenty-five test cases and threw a bunch of different LLMs at the problem. This is where the paper gets really practical about model choice.

Tom: And the results are fascinating. They compared a small local LLM against a larger one. The small one got a CFP accuracy of zero point four six, which means it struggled to even formulate the call for proposals correctly. The larger model got zero point five six.

Jane: But here's the twist. The larger model, when used in a single-round negotiation, only reached zero point seven one offloading accuracy. That means it picked the right destination less often than the small model did in multi-round mode, which hit zero point seven eight.

Tom: So model size alone doesn't win. The interaction structure matters more. When they let the larger model negotiate over multiple rounds, it jumped to zero point eight eight. That's a huge leap from zero point seven one.

Jane: It's like the difference between asking one person for directions and asking three people, then comparing their answers. The back-and-forth is what catches mistakes.

Tom: Exactly. And then they went further and tested five different LLMs with six different prompting strategies. Zero-shot, few-shot, chain-of-thought, ReAct, and combinations. Thirty configurations total.

Jane: And the winner? Gemini-three-Flash with a ReAct plus few-shot plus chain-of-thought hybrid hit a perfect one point zero zero offloading accuracy. But it wasn't the fastest or cheapest.

Tom: Right. And here's the warning in the data. For DeepSeek-V4-Pro, adding more prompting components actually hurt. It went from zero point seven two accuracy down to zero point six eight. More complexity isn't always better.

Jane: That's such an important finding. You can't just stack prompting techniques like Lego bricks. Each model has its own personality, and you have to find what works for that specific model.

Tom: And they also measured tokens and latency. The small local Llama3 model took one hundred eighty-six seconds per case. That's way too slow for real-time scheduling. The larger models were faster, around twenty to seventy-five seconds.

Jane: So there's a real trade-off between accuracy, speed, and cost. A perfect one point zero zero accuracy is great, but if it takes too long, the stream might miss its deadline anyway.

Tom: The authors also flagged that these results are descriptive, not definitive. Twenty-five cases is a small sample. One changed prediction shifts accuracy by four percentage points.

Jane: Still, the pattern is clear. You need to match the model and the prompt to the specific negotiation task. There's no one-size-fits-all.

Tom: And I think the biggest improvement this paper suggests is the idea of selective escalation. Don't run the LLM on every decision. Run it only when the deterministic rules hit a conflict or an ambiguity.

Jane: That's the smart engineering approach. Use the expensive, powerful tool only when you actually need it. That's how you get the best of both worlds.

Tom: And that's exactly where we're heading in the conclusion, because this has big implications for how we build real systems.

Conclusion — Tom and Jane summarize the paper 'Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing' and its implications, say goodbye to the paper and get ready to discuss the next one.: Jane: Alright, Tom, let's wrap this up. We've been talking about "Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing," and I think the core message is really elegant.

Tom: It is. The authors took a forty-year-old protocol, the Contract Net Protocol, and showed that the biggest gains come from just letting agents negotiate over multiple rounds. That alone cut latency violations in half and eliminated resource overcommitment.

Jane: And then the LLM adds the cherry on top. It handles the messy, qualitative context, the reliability warnings, the uncertain forecasts, and it pushes performance even higher.

Tom: But they were honest about the costs. The LLM is slower, it uses more tokens, and it's not always better. The best model-prompt combination hit perfect accuracy, but another model got worse when they added more prompting tricks.

Jane: So the real takeaway is about boundaries. Keep the hard constraints deterministic. Let the LLM handle interpretation and refinement. And escalate to the LLM selectively, only when the situation genuinely needs it.

Tom: And that's a blueprint we can actually build on. This isn't a paper that just says "AI is amazing." It says "here's exactly where AI helps, and here's where it doesn't."

Jane: I also appreciate that they released their benchmark, prompts, and raw outputs. That means other researchers can replicate and build on this work. That's how science moves forward.

Tom: For the real world, think about smart cities, autonomous vehicles, industrial IoT. Anywhere you have latency-sensitive streams running on distributed edge resources. This approach could make those systems more reliable without requiring a complete redesign.

Jane: And I think that's the most exciting part. You don't have to throw away your existing scheduling system. You just extend it with negotiation rounds and selectively add an LLM where it helps.

Tom: Well said, Jane. So we're saying goodbye to this paper, but we're taking its lessons with us. Next up, we've got a paper on federated learning for medical imaging that I think is going to spark some debate.

Jane: Can't wait. Thanks for listening, everyone. We'll see you on the next episode.

More episodes

← Home