An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing".
Jane: The paper was written by Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou and Malcolm Yoke Hean Low from Nanyang Technological University and Singapore Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. We’ve got a paper that’s got me genuinely fired up this morning. It’s called “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing.” That’s a mouthful, but the idea behind it is really cool.
Jane: It really is, Tom. And I think the title tells you exactly what’s going on. We’ve got drones, we’ve got AI, we’ve got factories. The authors are from Nanyang Technological University and the Singapore Institute of Technology, and they’re tackling this problem where drones aren’t just delivering products anymore. They’re also acting as flying computers.
Tom: Right, right. So in a smart factory, you have these stations that make stuff. Drones come by to pick up finished products. But at the same time, those stations have sensors generating computing tasks. The drone can help process those tasks while it’s there. That’s the mobile edge computing part.
Jane: Exactly. And the challenge is that the drone’s route decides when it’s at a station. That means the route decides when that station can offload its computing work to the drone. So you can’t plan the delivery route without thinking about the computing schedule, and you can’t plan the computing without knowing the route. They’re completely tangled up.
Tom: And that’s where the “agentic AI” part comes in. The authors built a system where you can just describe what you want in plain English, and the AI helps you write the mathematical model for this tangled problem. It’s like having a really smart assistant who knows operations research.
Jane: It’s a great example of using AI to help us build better AI systems. The paper is really about making this whole process accessible and reliable, which is a big deal for people who actually run these factories.
Tom: And they’re not just stopping at the model. They also built a way to solve it using reinforcement learning. So we’re talking about the full pipeline here, from describing the problem to getting a solution. I can’t wait to dig into how they actually did it.
Jane: Me neither. Let’s get into the details of the approach and the results.
Summary: Tom: So, Jane, we’ve got this paper, “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing.” Let’s break down what they actually did, because the summary is pretty dense.
Jane: It is dense, but the core idea is simple. They split the problem into two layers. The top layer is the drone routing. Which drone goes to which station, in what order, and when does it come back. The bottom layer is the task scheduling. For every second of the mission, which computing task gets processed, and where does it get processed.
Tom: And that split is smart because if you tried to solve the whole thing at once, the state space would be enormous. You’d have to think about every possible route and every possible computing decision at the same time. That’s just too much for a single learning agent.
Jane: Exactly. So they use a hierarchical approach. The upper layer learns the routing policy using a method called PPO, which is a popular reinforcement learning algorithm. Once that routing is fixed, the lower layer learns the scheduling policy, also using PPO. The routing layer passes down the service windows, meaning the times when each drone is at each station, and the lower layer uses that to decide if it can offload tasks.
Tom: And the results show this works really well. The upper layer, the routing, it learned to collect almost all the products. In the last five hundred training episodes, it achieved a ninety-nine point six percent collection rate. That’s nearly perfect.
Jane: That’s impressive. And the lower layer, the scheduling, it maintained a one hundred percent deadline satisfaction rate. Every single task finished on time. And they compared it to a different algorithm called A2C, and the PPO approach was much more stable in the later stages of training. A2C had these sudden drops in performance, but PPO just stayed solid.
Tom: Stability is huge in real-world applications. You don’t want a system that works great one day and then falls off a cliff the next. So the fact that they’ve got a method that’s both effective and stable is a really strong result.
Jane: And it’s all built on that agentic AI framework that helps you even formulate the problem in the first place. So the whole thing, from modeling to solving, is designed to be practical. I’m curious about how they actually built that AI assistant, though.
Tom: Good segue, because that’s exactly what we’re going to talk about next.
Improvements: Jane: So, Tom, we’ve talked about the two-layer reinforcement learning solution. But the part that really caught my eye in “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing” is the AI assistant that helps you write the math.
Tom: Yeah, that’s the part that feels like magic. You just tell the system what your factory looks like, and it writes down the equations for you. But how does it actually work without making stuff up?
Jane: That’s the key question. And the paper’s answer is a combination of two things. First, they use retrieval-augmented generation, or RAG. That means the AI doesn’t just rely on what it learned from the internet. It has a specific database of relevant knowledge about drone routing and task offloading, and it pulls from that database to ground its answers.
Tom: So it’s like giving the AI a textbook to look at instead of letting it guess from memory.
Jane: Exactly. And the second part is chain-of-thought reasoning. The AI doesn’t just spit out the whole model at once. It breaks the problem down into steps. First, define the objective. Then, write the routing constraints. Then, write the offloading constraints. Then, summarize the notation. Each step is checked by a separate verifier agent before the system moves on.
Tom: And that verification step is what makes it reliable. The verifier can say, “Hey, that constraint doesn’t match what’s in the database,” and the responder has to rethink and try again. It’s a back-and-forth conversation between two AI agents.
Jane: Right. And the paper shows that this process actually makes a difference. They compared their full framework to a simpler version that just had the responder without the chain-of-thought and verification. The simpler version changed the meaning of the objective function. It turned a penalty for “normalized resource occupation” into a penalty for “raw resource usage.” That’s a subtle but important difference.
Tom: That’s a great example. It sounds like the same thing, but it changes what the model is actually optimizing for. The full framework kept the original meaning intact because the verifier caught that drift.
Jane: So the improvement here isn’t just about making the AI faster or more accurate. It’s about making the AI more trustworthy. You can trace exactly why it made each decision, and you can catch mistakes before they become problems. That’s a huge step for using AI in industrial settings where errors are expensive.
Tom: And it means engineers who aren’t experts in operations research can still build these complex models. The AI is doing the heavy lifting, and the human is just guiding it. That could really change how factories are managed.
Jane: I think so. And it sets up a really interesting question about where this technology goes next. Let’s wrap up with our thoughts on that.
Conclusion: Tom: Alright, we’ve spent a lot of time with “An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing,” and I think we’re ready to say goodbye to it.
Jane: It’s been a great paper to talk about. We’ve got drones doing delivery and computing, an AI assistant that helps you write the math, and a two-layer reinforcement learning system that solves it all. The results were strong, with near-perfect collection rates and perfect deadline satisfaction.
Tom: And the real takeaway for me is that this isn’t just a theoretical exercise. The authors built a full pipeline. You describe your problem, the agentic AI helps you formulate it, and then the hierarchical DRL solves it. That’s a practical toolkit for real factories.
Jane: Absolutely. And I think the biggest impact will be on how we think about AI in manufacturing. It’s not just about automating a single task anymore. It’s about having AI that can understand the whole system, help you model it, and then optimize it. That’s a big shift.
Tom: And the future work is exciting too. The authors mention scaling to larger systems, heterogeneous drone fleets, and more dynamic environments. So this is really just the beginning.
Jane: I’m looking forward to seeing where it goes. Thanks for joining us, everyone. We’ll be back soon with another paper.
Tom: See you next time.
Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou, Malcolm Yoke Hean Low
Nanyang Technological University · Singapore Institute of Technology
cs.AI, cs.LG
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 15 pages
Code: https://github.com/Puppet88/Agentic-AI-UAV
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
Key concepts
- UAV-Assisted Logistics Scheduling
- This refers to planning drone routes (UAVs) in a factory setting. The drones must balance picking up finished products while also determining when and where they can offload computing tasks from various stations along their route.
- Mobile Edge Computing
- In this context, it means the drone acts as a mobile computer. While moving through the factory, it processes computing tasks generated by stationary sensors and stations, utilizing its own onboard resources.
- Agentic AI Framework
- This is an AI system that allows users to describe complex problems in plain English. The framework then helps generate the necessary mathematical model for solving the problem, making advanced operations research accessible to non-experts.
- Chain-of-Thought (CoT) Reasoning
- This technique improves AI reliability by forcing the model to break down a problem into sequential steps. Each step is checked by a verifier agent, ensuring accuracy and preventing subtle errors when formulating complex mathematical models.
Terminology
Summary
Summary
This paper studies a hybrid scheduling problem in cloud manufacturing (CMfg) where unmanned aerial vehicles (UAVs) support both product collection and mobile edge computing (MEC). The authors state: "In this paper, UAVs collect finished products from manufacturing stations and transport them back to a central depot. Meanwhile, computational tasks generated by industrial sensor devices at these stations are processed locally, at UAVs, or offloaded via UAVs to the cloud. The coupling between physical logistics and computational task scheduling makes the problem challenging because
A UAV can provide MEC services only during its service window at a station, so routing decisions directly determine when UAV-assisted offloading is available. Routing decisions also affect the UAV energy budget and the availability of onboard computing and communication resources for computational task execution under task deadline constraints."
The paper identifies two main challenges. Challenge I is that The mathematical modeling of combinatorial optimization problems in CMfg can be significantly more complex than that of traditional single-domain scheduling problems,
particularly because "It requires capturing how UAV routing induces time-varying service windows. These service windows, in turn, affect per-slot task offloading and compute/communication allocation under energy budgets, capacity limits, and task deadlines. Challenge II is that
Compared with traditional scheduling problems, hybrid task scheduling in CMfg introduces significantly higher solution complexity because
logistics- and computation-related decisions must be optimized jointly rather than handled in isolation."
To address Challenge I, the authors propose "an agentic AI framework to address Challenge I. The agent facilitates dialogue-driven mathematical model generation and refinement by combining LLM reasoning with retrieval augmented generation (RAG) and chain-of-thought (CoT) reasoning. Specifically,
an LLM converts natural language descriptions into structured modeling logic. Then, RAG extracts pertinent formulations, constraints, and objectives from curated literature and technical documentation. CoT further makes explicit the logic underlying variable definitions, constraint design, and objective construction."
To address Challenge II, the authors propose a hierarchical DRL approach to solve the coupled UAV-assisted collection and mobile edge computing (MEC) scheduling problem.
The framework "decomposes the original joint optimization problem into two tractable sequential decision layers, which substantially reduces the effective state/action space while preserving the essential coupling between logistics and computation services. Both layers are
formulated as finite-horizon Markov decision processes (MDPs). They are solved using proximal policy optimization (PPO). The upper layer
formulates the multi-UAV routing problem as a single-agent MDP. It jointly determines the manufacturing-station assignment for each UAV and the corresponding visiting order under energy, payload, and flight-distance constraints. The lower layer
makes per-slot execution decisions—offloading destination (local/UAV/cloud via UAV) and compute/communication allocation—under capacity and deadline constraints. The two layers interact through two signals:
The upper layer provides service windows and remaining UAV energy to the lower layer, which uses them to determine the feasibility of UAV-involved processing."
The system model considers a fleet of U homogeneous UAVs dispatched from a central depot with terrestrial MEC servers (referred to as the cloud) to collect finished products from M geographically distributed manufacturing stations. Each station has a collection reward, product weight, and hosts a set of industrial sensor devices (ISDs) that generate computational tasks stochastically over a discretized time horizon. The UAV routing model includes constraints on collection assignment and route consistency, payload and distance budgets, and mission timing feasibility. The MEC service model includes constraints on execution-mode selection, completion-time encoding, computing and communication resource limits, service-window indicators, service-window feasibility for UAV-involved processing, deadline satisfaction, and UAV energy budgeting.
The complete optimization problem is formulated as: P: max ωcol Σ vm cm + ωcmp Σ zk,τ − ωmiss Σ (1 − zk,τ) − ωflow Σ (Tk,τ − τ∆) − ωres Σ (Rtcmp + Rtcom), s.t. Eq. (1) − Eq. (9), Eq. (10) − Eq. (27).
The objective promotes product collection and deadline-compliant task completion, while reducing deadline violations, task completion delay, and overall computing and communication resource occupation.
The problem analysis shows that "the UAV routing part is NP-hard since it jointly determines station assignment and visit sequencing under route-consistency, payload, distance, and mission time constraints while maximizing collection value, which amounts to a capacitated value-collecting multi-UAV routing variant. The task offloading part is also NP-hard. Therefore,
both subproblems, and thus the full joint formulation, are NP-hard, which motivates our proposed hierarchical DRL approach."
The proposed agentic AI framework includes a RAG module that is used to ground the LLM with task-relevant evidence before model generation.
The RAG module organizes an internal knowledge base from two perspectives: UAV routing and computational task offloading.
The retrieval process uses cosine similarity to measure relevance between the query and each chunk, and the top-Kret most relevant chunks form the context for generation. The CoT module reasons over the retrieved information and organizes it for downstream mathematical formulation.
The framework includes two agents: "the Agentic AI Responder is responsible for reasoning over the user request and providing the final answer. In contrast, the Agentic AI Verifier evaluates another agent's responses, deciding whether to approve or reject it. The CoT workflow decomposes the user request into four steps:
objective function formulation, UAV routing constraints derivation, task offloading constraints derivation, and mathematical notation summarization. For each CoT step,
the Agentic AI Responder initially formulates a response to the specific CoT-step query; subsequently, the Agentic AI Verifier assesses the generated response prior to the continuation of the reasoning process."
The hierarchical DRL approach uses PPO in both layers. The upper-layer MDP state "summarizes four types of information: (i) per-UAV status, including location, remaining energy, payload, and flight-distance budget; (ii) route-progress information; (iii) per-station attributes; and (iv) mission-level context. The upper-layer action includes
a priority score and
routing actions, where each UAV selects
one discrete routing symbol from stay, add(1),..., add(M), complete. The upper-layer reward includes
collection value obtained during the transition," a shaping reward that encourages increasing the coverage ratio of collected stations,
balanced route workloads across UAVs,
and route termination
rewards, minus an aggregated penalty that discourages infeasible decisions.
The lower-layer MDP state "summarizes blocal queues, UAV-side resources, and current service feasibility through six groups of normalized features: (i) per-ISD local statistics; (ii) per-UAV statistics; (iii) global system scalars; (iv) per UAV–ISD active offload status; (v) a UAV–ISD service-window availability mask; and (vi) a fixed-length, zero-padded snapshot of the top-Nq unfinished tasks. The lower-layer action is
defined over a cached global-queue snapshot Qtb and specifies
the service decisions for the first Kg queue positions. Each action tuple includes
a processing choice for the queue item at position j selected from
SKIP, local ∪ UAV(u), cloud-via-UAV(u) U u=1 and
normalized compute and communication allocations that are
quantized by the environment into discrete levels. The lower-layer reward
combines positive incentives, penalties, and an end-of-episode settlement," including rewards for task completions, dispatch bonuses, dense progress rewards, shaping terms for backlog and urgency reduction, and penalties for deadline violations, infeasible decisions, avoidable idling, excessive resource usage, and remaining unfinished load.
The complexity analysis shows that the upper-layer training complexity is O(Nup T̄ up (U M + Pθ,ϕ up))
and the lower-layer training complexity is O(Nlo T̄ lo (U Kg + Nq + Kg U + Pθ,ϕ lo)).
The authors note that "a monolithic MDP would need to enumerate joint routing and task-scheduling choices, whose size scales with (M + 2)U (2U + 2)Kg before considering resource-allocation levels. Therefore, the proposed hierarchical decomposition keeps the rollout and update costs polynomial in the main system sizes while preserving the service-window coupling between UAV routing and MEC task scheduling."
The training procedure is sequential: "First, the upper layer is trained by PPO for UAV routing... The best-performing upper-layer policy is then fixed to extract the UAV service-window and residual-energy information for the lower layer. Next, the lower layer is trained by PPO for task scheduling and resource allocation under this fixed routing information."
Simulation results are presented for both the agentic AI framework and the hierarchical DRL approach. For the agentic AI framework, the results show that "the proposed framework remains closer to the original modeling logic by preserving the intended meanings of the modeled terms, whereas the baseline tends to reformulate them into a different optimization-oriented representation. For example,
In the RAG database, the penalty is defined through the abstract occupation variables Rtcmp and Rtcom, which represent normalized computing and communication occupation. The proposed framework preserves this same interpretation... By contrast, the baseline rewrites this penalty as a direct sum of computing and communication allocation variables. As a result, it changes the penalty from normalized resource occupation to raw resource usage."
For the upper-layer DRL training, "the upper-layer PPO shows good learning ability and convergence. In the early stage, the Raw reward fluctuates sharply, indicating active exploration. With training, MA(50) increases rapidly and then stabilizes around 400, indicating that the policy is improved rapidly and converges to a stable high reward solution. The collection rate
quickly approaches 100% and remains at or near 100% in most episodes, indicating that the PPO policy can reliably plan UAV routing and complete almost all collection tasks."
For the lower-layer DRL training, "both PPO and A2C improve the lower-layer policy from highly negative rewards to a converged positive-reward region. PPO exhibits larger fluctuations in the intermediate stage, but its reward quickly recovers and remains relatively stable afterward. In contrast, the total reward under A2C increases more smoothly in the early stage, but several sharp reward drops still appear in the later stage, indicating weaker stability after convergence. Regarding deadline satisfaction,
PPO maintains a rate very close to 100% more consistently in the later stage, whereas A2C still experiences several late-stage drops and cannot always sustain the ideal 100% deadline satisfaction rate."
The paper concludes that "the learned policies can achieve strong routing and scheduling performance, while the overall design preserved the essential coupling between logistics and computation through compact cross-layer information exchange. Overall, this work provides a unified framework for both formulating and solving hybrid logistics-computation scheduling problems in CMfg. Future work can extend the framework to larger-scale systems, heterogeneous UAV fleets, and more dynamic manufacturing environments."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
Improvement: Implement a two-agent (Responder + Verifier) architecture with Chain-of-Thought (CoT) decomposition and Retrieval-Augmented Generation (RAG) for automated optimization problem formulation.
Capabilities:
-
Parse natural-language system descriptions into structured mathematical models (objective functions, constraints, decision variables)
-
Decompose complex formulation tasks into 4 sequential CoT steps: objective definition, routing constraints, offloading constraints, notation summarization
-
Retrieve domain-specific modeling knowledge (UAV routing, MEC offloading) from a curated vector database using cosine similarity
-
Verify each intermediate reasoning step before proceeding, with automatic rethinking on failure
-
Produce interpretable, traceable formulations that preserve semantic fidelity to domain knowledge (unlike single-agent baselines that alter meaning)
These improvements enable AI systems to: automatically generate correct optimization models from natural language, solve coupled logistics-computation problems with high reliability, maintain deadline guarantees in dynamic environments, and scale to larger problem instances with polynomial training complexity.
Abstract
In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational task scheduling. In this paper, UAVs collect finished products from manufacturing stations and transport them back to a central depot. Meanwhile, computational tasks generated by industrial sensor devices at these stations are processed locally, at UAVs, or offloaded via UAVs to the cloud. This coupling makes the problem challenging. A UAV can provide MEC services only during its service window at a station, so routing decisions directly determine when UAV-assisted offloading is available. Routing decisions also affect the UAV energy budget and the availability of onboard computing and communication resources for computational task execution under task deadline constraints. To address this, we propose an agentic-AI-assisted optimization framework with two components. First, we develop an agentic AI that combines large language models, retrieval-augmented generation, and chain-of-thought reasoning to translate user input into an interpretable mathematical formulation for the hybrid scheduling problem. Second, we design a hierarchical deep reinforcement learning approach based on proximal policy optimization (PPO), where the upper layer learns UAV routing and the lower layer optimizes per-slot task execution and resource allocation. Simulation results show that the proposed framework yields more consistent formulations, while the hierarchical PPO achieves full product collection in 99.6% of the last 500 episodes and maintains a 100% deadline satisfaction rate, with more stable performance than the advantage actor-critic approach.
Sources
- GraphThought: Graph Combinatorial Optimization with Thought Generation
- LLMs can Schedule
- A Large Language Model-based multi-agent manufacturing system for intelligent shopfloor
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
- Zero-Shot Verification-guided Chain of Thoughts
- Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
- Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection