An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing

summary

Video file (mp4)

In short

The episode discusses 'An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing.' Hosts review a system that uses drones for both product delivery and computing tasks in smart factories. They detail the two-layer reinforcement learning approach and the agentic AI framework used to model complex problems.

Key concepts

UAV-Assisted Logistics Scheduling
This refers to planning drone routes (UAVs) in a factory setting. The drones must balance picking up finished products while also determining when and where they can offload computing tasks from various stations along their route.
Mobile Edge Computing
In this context, it means the drone acts as a mobile computer. While moving through the factory, it processes computing tasks generated by stationary sensors and stations, utilizing its own onboard resources.
Agentic AI Framework
This is an AI system that allows users to describe complex problems in plain English. The framework then helps generate the necessary mathematical model for solving the problem, making advanced operations research accessible to non-experts.
Chain-of-Thought (CoT) Reasoning
This technique improves AI reliability by forcing the model to break down a problem into sequential steps. Each step is checked by a verifier agent, ensuring accuracy and preventing subtle errors when formulating complex mathematical models.

Terminology used across episodes

This episode discusses

The paper

An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing · Read on arXiv

Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou, Malcolm Yoke Hean Low

Nanyang Technological University · Singapore Institute of Technology

In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational task scheduling. In this paper, UAVs collect finished products from manufacturing stations and transport them back to a central depot. Meanwhile, computational tasks generated by industrial sensor devices at these stations are processed locally, at UAVs, or offloaded via UAVs to the cloud. This coupling makes the problem challenging. A UAV can provide MEC services only during its service window at a station, so routing decisions directly determine when UAV-assisted offloading is available. Routing decisions also affect the UAV energy budget and the availability of onboard computing and communication resources for computational task execution under task deadline constraints. To address this, we propose an agentic-AI-assisted optimization framework with two components. First, we develop an agentic AI that combines large language models, retrieval-augmented generation, and chain-of-thought reasoning to translate user input into an interpretable mathematical formulation for the hybrid scheduling problem. Second, we design a hierarchical deep reinforcement learning approach based on proximal policy optimization (PPO), where the upper layer learns UAV routing and the lower layer optimizes per-slot task execution and resource allocation. Simulation results show that the proposed framework yields more consistent formulations, while the hierarchical PPO achieves full product collection in 99.6% of the last 500 episodes and maintains a 100% deadline satisfaction rate, with more stable performance than the advantage actor-critic approach.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing".

Jane: The paper was written by Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou and Malcolm Yoke Hean Low from Nanyang Technological University and Singapore Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. We’ve got a paper that’s got me genuinely fired up this morning. It’s called “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing.” That’s a mouthful, but the idea behind it is really cool.

Jane: It really is, Tom. And I think the title tells you exactly what’s going on. We’ve got drones, we’ve got AI, we’ve got factories. The authors are from Nanyang Technological University and the Singapore Institute of Technology, and they’re tackling this problem where drones aren’t just delivering products anymore. They’re also acting as flying computers.

Tom: Right, right. So in a smart factory, you have these stations that make stuff. Drones come by to pick up finished products. But at the same time, those stations have sensors generating computing tasks. The drone can help process those tasks while it’s there. That’s the mobile edge computing part.

Jane: Exactly. And the challenge is that the drone’s route decides when it’s at a station. That means the route decides when that station can offload its computing work to the drone. So you can’t plan the delivery route without thinking about the computing schedule, and you can’t plan the computing without knowing the route. They’re completely tangled up.

Tom: And that’s where the “agentic AI” part comes in. The authors built a system where you can just describe what you want in plain English, and the AI helps you write the mathematical model for this tangled problem. It’s like having a really smart assistant who knows operations research.

Jane: It’s a great example of using AI to help us build better AI systems. The paper is really about making this whole process accessible and reliable, which is a big deal for people who actually run these factories.

Tom: And they’re not just stopping at the model. They also built a way to solve it using reinforcement learning. So we’re talking about the full pipeline here, from describing the problem to getting a solution. I can’t wait to dig into how they actually did it.

Jane: Me neither. Let’s get into the details of the approach and the results.

Summary: Tom: So, Jane, we’ve got this paper, “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing.” Let’s break down what they actually did, because the summary is pretty dense.

Jane: It is dense, but the core idea is simple. They split the problem into two layers. The top layer is the drone routing. Which drone goes to which station, in what order, and when does it come back. The bottom layer is the task scheduling. For every second of the mission, which computing task gets processed, and where does it get processed.

Tom: And that split is smart because if you tried to solve the whole thing at once, the state space would be enormous. You’d have to think about every possible route and every possible computing decision at the same time. That’s just too much for a single learning agent.

Jane: Exactly. So they use a hierarchical approach. The upper layer learns the routing policy using a method called PPO, which is a popular reinforcement learning algorithm. Once that routing is fixed, the lower layer learns the scheduling policy, also using PPO. The routing layer passes down the service windows, meaning the times when each drone is at each station, and the lower layer uses that to decide if it can offload tasks.

Tom: And the results show this works really well. The upper layer, the routing, it learned to collect almost all the products. In the last five hundred training episodes, it achieved a ninety-nine point six percent collection rate. That’s nearly perfect.

Jane: That’s impressive. And the lower layer, the scheduling, it maintained a one hundred percent deadline satisfaction rate. Every single task finished on time. And they compared it to a different algorithm called A2C, and the PPO approach was much more stable in the later stages of training. A2C had these sudden drops in performance, but PPO just stayed solid.

Tom: Stability is huge in real-world applications. You don’t want a system that works great one day and then falls off a cliff the next. So the fact that they’ve got a method that’s both effective and stable is a really strong result.

Jane: And it’s all built on that agentic AI framework that helps you even formulate the problem in the first place. So the whole thing, from modeling to solving, is designed to be practical. I’m curious about how they actually built that AI assistant, though.

Tom: Good segue, because that’s exactly what we’re going to talk about next.

Improvements: Jane: So, Tom, we’ve talked about the two-layer reinforcement learning solution. But the part that really caught my eye in “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing” is the AI assistant that helps you write the math.

Tom: Yeah, that’s the part that feels like magic. You just tell the system what your factory looks like, and it writes down the equations for you. But how does it actually work without making stuff up?

Jane: That’s the key question. And the paper’s answer is a combination of two things. First, they use retrieval-augmented generation, or RAG. That means the AI doesn’t just rely on what it learned from the internet. It has a specific database of relevant knowledge about drone routing and task offloading, and it pulls from that database to ground its answers.

Tom: So it’s like giving the AI a textbook to look at instead of letting it guess from memory.

Jane: Exactly. And the second part is chain-of-thought reasoning. The AI doesn’t just spit out the whole model at once. It breaks the problem down into steps. First, define the objective. Then, write the routing constraints. Then, write the offloading constraints. Then, summarize the notation. Each step is checked by a separate verifier agent before the system moves on.

Tom: And that verification step is what makes it reliable. The verifier can say, “Hey, that constraint doesn’t match what’s in the database,” and the responder has to rethink and try again. It’s a back-and-forth conversation between two AI agents.

Jane: Right. And the paper shows that this process actually makes a difference. They compared their full framework to a simpler version that just had the responder without the chain-of-thought and verification. The simpler version changed the meaning of the objective function. It turned a penalty for “normalized resource occupation” into a penalty for “raw resource usage.” That’s a subtle but important difference.

Tom: That’s a great example. It sounds like the same thing, but it changes what the model is actually optimizing for. The full framework kept the original meaning intact because the verifier caught that drift.

Jane: So the improvement here isn’t just about making the AI faster or more accurate. It’s about making the AI more trustworthy. You can trace exactly why it made each decision, and you can catch mistakes before they become problems. That’s a huge step for using AI in industrial settings where errors are expensive.

Tom: And it means engineers who aren’t experts in operations research can still build these complex models. The AI is doing the heavy lifting, and the human is just guiding it. That could really change how factories are managed.

Jane: I think so. And it sets up a really interesting question about where this technology goes next. Let’s wrap up with our thoughts on that.

Conclusion: Tom: Alright, we’ve spent a lot of time with “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing,” and I think we’re ready to say goodbye to it.

Jane: It’s been a great paper to talk about. We’ve got drones doing delivery and computing, an AI assistant that helps you write the math, and a two-layer reinforcement learning system that solves it all. The results were strong, with near-perfect collection rates and perfect deadline satisfaction.

Tom: And the real takeaway for me is that this isn’t just a theoretical exercise. The authors built a full pipeline. You describe your problem, the agentic AI helps you formulate it, and then the hierarchical DRL solves it. That’s a practical toolkit for real factories.

Jane: Absolutely. And I think the biggest impact will be on how we think about AI in manufacturing. It’s not just about automating a single task anymore. It’s about having AI that can understand the whole system, help you model it, and then optimize it. That’s a big shift.

Tom: And the future work is exciting too. The authors mention scaling to larger systems, heterogeneous drone fleets, and more dynamic environments. So this is really just the beginning.

Jane: I’m looking forward to seeing where it goes. Thanks for joining us, everyone. We’ll be back soon with another paper.

Tom: See you next time.

More episodes

← Home