E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments".
Jane: The paper was written by Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha et al. from University of Oslo and UiT The Arctic University of Norway.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone! Today we're diving into a fresh arXiv paper that's got me genuinely pumped. It's called "E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments." Jane, what caught your eye first?
Jane: Oh, Tom, the title alone tells you we're not in cloud territory anymore. We're talking about running these massive language models on the devices we actually have around us — your laptop, a Jetson board, maybe even something smaller. The authors are from the University of Oslo and UiT The Arctic University of Norway, and they're tackling a problem that's been nagging at me for a while.
Tom: Right? Most people think LLMs just live in the cloud, but the paper points out that a lot of organizations — small companies, clinics, even individual developers — want to keep their data local. They don't want to ship sensitive information to a cloud provider.
Jane: Exactly. And the authors make this great point early on: the old assumption that one device can hold the entire model just doesn't hold up in edge environments. A single Raspberry Pi or even a decent workstation can't store a twenty-billion-parameter model in memory. So they're asking, how do we split this thing up and still make it fast?
Tom: And that's where the name comes from — E2LLM. It's about efficiency, but also about being practical. The team isn't just theorizing; they actually built a framework and tested it on real hardware.
Jane: I love that they're not pretending the cloud doesn't exist. They're saying, look, if you're in a resource-limited setting, you need a different playbook. And that playbook involves replication — running multiple copies of the model across groups of devices — plus something called model parallelism, which is basically splitting the model's layers across several machines.
Tom: So instead of one giant pipeline with every device in the network, they form smaller clusters, each hosting its own full copy of the model. That's a clever twist on the usual approach.
Jane: It is. And the authors are upfront that finding the best way to group devices is a hard problem — NP-hard, actually. So they use a Genetic Algorithm to search for good solutions. We'll get into that in a bit, but the headline is: they're not just throwing hardware at the problem; they're being smart about how to organize it.
Tom: And the implications? If this works, it means we could see LLMs running in places we never thought possible — hospitals, small businesses, even community networks. That's a big deal for privacy and for accessibility.
Jane: Absolutely. And the paper's results suggest they're onto something. But before we get into the numbers, let's talk about what they actually built. That's coming up next.
Paper Summary: Tom: So we've got the title and the big idea — running LLMs on edge and fog devices. But what does E2LLM actually do? Jane, walk us through the core mechanism.
Jane: Okay, so imagine you have a bunch of devices — some are beefy, some are weak, all connected over a local network. The first thing E2LLM does is profile each device: how fast can it process a layer of the model? How much memory does it have? What's the network bandwidth between devices?
Tom: So it's like taking inventory before you plan the party.
Jane: Exactly. Then it uses a Genetic Algorithm to group these devices into clusters. Each cluster becomes a replica of the full model. But here's the twist — not all replicas do the same job. The system separates the two phases of LLM inference: the prefill phase, where the model reads all the input tokens at once, and the decode phase, where it generates one token at a time.
Tom: And that separation matters because those phases have very different characteristics, right?
Jane: Right. Prefill is compute-bound — it's doing a ton of parallel math on all the input tokens at once. Decode is memory-bound — it's mostly moving data around, generating one token at a time. So E2LLM assigns some replicas to be prefill specialists and others to be decoder specialists. It's like having a chef who's great at prepping ingredients and another who's great at plating.
Tom: That's a smart division of labor. And the paper mentions Splitwise as a baseline — that's an existing system that also splits prefill and decode. But E2LLM goes further.
Jane: It does. Splitwise assumes each device can hold the whole model. E2LLM doesn't make that assumption. Instead, within each cluster, it uses dynamic programming to figure out how to split the model's layers across the devices in that cluster. The goal is to minimize the bottleneck — the slowest stage in the pipeline.
Tom: And why is that bottleneck so important?
Jane: Because in a pipeline, the slowest stage sets the pace for everything. If one device is struggling, the whole cluster waits. So E2LLM's dynamic programming finds a partitioning that balances the load as evenly as possible, while respecting memory constraints.
Tom: So it's not just about grouping devices — it's about making each group work as efficiently as possible.
Jane: Precisely. And then, once they have the deployment plan, they use a load balancer to distribute incoming requests. The paper shows that this approach can handle high-demand scenarios much better than the baseline.
Tom: And the results? I saw something about cutting waiting time by more than half.
Jane: Yeah, in high-demand conditions, E2LLM reduced average waiting time by over fifty percent compared to Splitwise. And decoding throughput was up to twice as high. That's a huge improvement for user experience.
Tom: That's the kind of result that makes you sit up and take notice. But how do they actually get there? Let's dig into the methodology next.
Improvements Suggested: Tom: We've covered the big picture, but I want to get into the weeds a bit. What are the specific improvements E2LLM brings over existing methods?
Jane: Great question. The first improvement is the replication strategy itself. Most model parallelism approaches assume all devices form one giant pipeline. E2LLM says, no — let's form multiple smaller pipelines, each hosting a full copy of the model. That way, you get more parallelism at the request level, not just within a single inference.
Tom: So instead of one assembly line, you have several assembly lines running in parallel.
Jane: Exactly. And the second improvement is the role differentiation. The paper recognizes that prefill and decode have different computational profiles. Prefill is all about processing a big batch of input tokens quickly. Decode is about generating tokens one at a time, which is memory-bound. So E2LLM lets each replica specialize.
Tom: And that's different from Splitwise how?
Jane: Splitwise does the same phase separation, but it assumes a single device can host the whole model. E2LLM doesn't have that luxury. So it combines phase separation with model parallelism — you get the benefits of both.
Tom: And then there's the dynamic programming for partitioning. Can you explain why that's such an improvement?
Jane: Sure. The paper's dynamic programming approach explicitly minimizes the bottleneck stage. That's different from some other methods that just try to balance total computation time. The authors point out that when you have multiple batches flowing through a pipeline, the slowest stage determines the throughput. So they optimize for that directly.
Tom: And the Genetic Algorithm — what's that doing that a simple greedy approach couldn't?
Jane: The search space is enormous. You have to decide which devices go in which cluster, and in what order. The Genetic Algorithm explores that space intelligently, using crossover and mutation to evolve good solutions. It's not guaranteed to find the absolute best, but it finds good ones fast.
Tom: And they also consider the ratio of input tokens to generated tokens. That's interesting — they actually analyzed a dataset to see what that ratio looks like in practice.
Jane: Right. They looked at the LongReason dataset and found that the ratio varies a lot depending on the task. That ratio helps decide how many prefill replicas you need versus decoder replicas. If requests have lots of input tokens, you need more prefill capacity.
Tom: So it's adaptive — the deployment plan changes based on the expected workload.
Jane: Exactly. And that's a big deal. The system isn't just a one-size-fits-all deployment. It's tuned to the actual demands it'll face.
Tom: Now, what does this mean for real-world applications? Let's bring in Lu and Meng to get their takes.
Lu: I'm excited about the privacy angle. If you can run a decent LLM on your own edge infrastructure, you don't have to send sensitive data to a cloud provider. That's huge for healthcare, legal, finance — any field where data confidentiality matters.
Meng: And from an engineering standpoint, the fact that they tested on real hardware — seven different devices, from an RTX five thousand seventy to a Jetson AGX Orin — that's reassuring. It's not just a simulation. They actually deployed Gpt-Oss-20b and measured real throughput and latency.
Jane: And the numbers back it up. In high-demand scenarios, E2LLM cut waiting time by more than half. That's the kind of improvement that makes a system feel responsive instead of sluggish.
Tom: So the improvements aren't just theoretical — they're measurable and practical. But what's the bigger picture? What does this mean for the world?
Conclusion: Tom: We've covered a lot of ground on "E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments." Let's wrap this up. Jane, what's the one thing you want listeners to remember?
Jane: I think it's that you don't need a cloud data center to run a useful LLM. E2LLM shows that with smart clustering, role specialization, and careful partitioning, you can get solid performance from a handful of ordinary devices. That opens the door to local, private, cost-effective AI.
Tom: And the numbers back that up — up to twice the decoding throughput and more than fifty percent less waiting time compared to Splitwise. That's not a small gain.
Jane: Right. And the authors were thoughtful about the design. They profiled each device, used a Genetic Algorithm to find good clusterings, and used dynamic programming to minimize bottlenecks. It's a well-rounded approach.
Lu: I'd add that the implications go beyond just technical performance. This could democratize access to LLMs. Small clinics, rural schools, community organizations — they could run their own models without relying on big tech clouds. That's a meaningful step toward more equitable AI access.
Meng: And from a practical standpoint, the fact that they validated on real hardware with a real model makes me confident this could be adopted in production. The load balancing strategy — Join Shortest Queue — is simple and effective.
Tom: So what's the future work? The paper mentions network-aware optimization and dynamic scaling as next steps. That makes sense — if you can adapt to changing network conditions or workload patterns in real time, you could squeeze out even more performance.
Jane: Absolutely. And the authors acknowledge that their current approach assumes a relatively stable environment. Making it adaptive would be the next frontier.
Tom: Well, I think we've given "E2LLM" a proper send-off. It's a solid piece of research with real-world potential. Thanks to everyone who tuned in — we'll be back with another paper soon. Until then, keep exploring.
Jane: And if you're thinking about deploying an LLM on your own hardware, this paper is a great starting point. See you next time!
Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan
University of Oslo · UiT The Arctic University of Norway
cs.DC, cs.AI
Submitted: 2026-08-23
Updated: 2026-08-25
Code: https://github.com/ggml-org/llama.cpp
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 78/100
Key concepts
- E2LLM
- A framework designed for efficient LLM serving in edge/fog environments. It uses device profiling, a Genetic Algorithm for clustering, and dynamic programming to partition the model layers across devices to minimize bottlenecks.
- Model Parallelism
- Splitting a large language model's layers across several machines. E2LLM uses this technique within clusters to distribute the model's workload efficiently among participating devices.
- Prefill Phase
- The initial stage of LLM inference where the model reads all input tokens at once. This phase is compute-bound, requiring parallel processing of a large batch of input tokens.
- Decode Phase
- The subsequent stage where the model generates one token at a time. This phase is memory-bound, focusing on moving data around to produce sequential output.
Terminology
Summary
Summary
This paper introduces E2LLM, a framework for efficient Large Language Model (LLM) deployment in heterogeneous Edge and Fog environments, where individual devices lack the resources to host an entire model. The authors argue that conventional deployment strategies assume a single device can host the full model, which is impractical for edge/fog scenarios. They also note that existing model parallelism approaches typically assume all devices participate in a single pipeline, which increases network demand and can lead to resource under-utilization. E2LLM addresses these limitations by replicating the full model across multiple groups of devices (replicas) and applying model parallelism within each replica.
The core architecture of E2LLM involves partitioning available devices into multiple clusters, each hosting one replica of the target model. Each replica is assigned a specialized role—either PREFILL or DECODER—based on its efficiency in handling input and output tokens. This separation leverages the inherent differences between the two phases of LLM inference: the Prefill phase is generally compute-bound, while the Decoder phase tends to be memory-bound.
The authors use a Genetic Algorithm (GA) to form clusters that maximize system performance, and within each cluster, they apply Dynamic Programming to determine an optimal partitioning strategy that minimizes bottlenecks in model-parallel execution.
The paper's methodology consists of several key components. First, latency profiling is performed at the layer level, assuming that Transformer blocks of the same type have similar latency. The authors adapted llama.cpp to load the model with a reduced number of Transformer blocks to measure latency efficiently. Second, a Dynamic Programming algorithm (Algorithm 1) is proposed for model partitioning in pipeline parallelism, with the objective of minimizing the slowest stage in the pipeline. The algorithm has a complexity of O(M2 × N2), which is much lower than O(M2 × N2 × 2 M) of EdgeShard.
Third, the authors analyze Prefill and Decoder phases to determine role assignment criteria, using the bottleneck phase formula: bottleneck phase = max(NP/PS, ND/DS), where NP and ND are average input and generated tokens, and PS and DS are Prefill and Decoder speeds. Fourth, a two-chromosome Genetic Algorithm (Algorithm 2) is designed for deployment planning, where the first chromosome encodes node ordering and the second specifies grouping into replicas. The GA incorporates an elite preservation strategy, crossover, and mutation policies.
For Quality-of-Service (QoS), the authors focus on text generation speed, citing that the average human reading speed is about 238 words per minutes, roughly four words per second
and that fast readers can process between 7 and 15 words per second. Therefore, the system should generate text at a minimum rate of 7 tokens per second.
The experimental setup uses seven heterogeneous devices (Table II), ranging from embedded to desktop devices, including RTX 5070, Apple M1, Apple M2 Max, RTX 3060M, and Jetson AGX Orin, all connected via a high-speed LAN with 920Mbps bandwidth. The target model is Gpt-Oss-20b with 24 Transformer blocks. The dataset used is lz1bytedance/LongReason
in two versions: extended (average 576 input tokens, 588 generated tokens, ratio 0.98) and custom extended (average 2284 input tokens, 1004 generated tokens, ratio 2.27).
The authors compare E2LLM against an adapted SplitWise baseline, excluding HexGen and ThunderServe due to their optimization objectives not being applicable in this setting. The deployment plans (Tables III-VI) show that E2LLM can process 22 decoder requests concurrently compared to 17 for SplitWise with the extended dataset, and 18 versus 16 for the custom extended dataset.
Experimental results demonstrate that E2LLM consistently outperforms SplitWise across all arrival periods (0.5s, 1.0s, 2.0s, 3.0s) and both datasets. Key findings include: E2LLM achieves up to 2× higher decoding throughput (e.g., 30.0 tokens/s versus 12.1 tokens/s for SplitWise at 3.0s arrival period with extended dataset), and reduces maximum waiting time by more than half in high-demand conditions (e.g., 75.1s versus 191.1s at 0.5s arrival period with extended dataset). The paper states: Compared to the Splitwise baseline, E2LLM reduces average waiting time by over 50% under high-demand conditions.
In low-demand scenarios, E2LLM achieves waiting times of about 3.5 seconds including KV cache transmission time, while SplitWise exhibits values of about 17.1 seconds.
The authors conclude that E2LLM not only scales better under heavy load but also provides more predictable and lower latency in resource-constrained Edge/Fog environments.
They note that SplitWise performs poorly in these conditions because its design does not account for the timing-sensitive nature of edge workloads,
as it was originally designed for scenarios where timing is less critical and energy consumption is the primary optimization goal.
Future work will focus on network-aware optimization and dynamic scaling.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
-
Improvement: Replace the assumption that a full LLM must reside on a single device with a dynamic clustering algorithm that partitions the model across multiple heterogeneous devices (e.g., mixing RTX 5070, Apple M1, Jetson AGX Orin) based on real-time memory, compute, and network bandwidth profiling.
-
What it can do: Deploy a 20B-parameter LLM on a cluster of 7 mixed consumer/edge devices (total VRAM 100GB) without cloud dependency, achieving near-cloud performance.
-
Improvement: Automatically classify each replica as either a Prefill (compute-bound, processes input tokens) or Decoder (memory-bound, generates output tokens) based on the measured input-to-output token ratio of the actual workload, rather than a fixed manual assignment.
-
What it can do: In a workload with 2.27× more input tokens than output (custom dataset), the system automatically allocates more compute-heavy devices to Prefill and memory-heavy devices to Decoder, reducing the bottleneck by 50% compared to a uniform assignment.
-
Improvement: Replace exhaustive search (O(M2 × N2 × 2 M)) with a dynamic programming algorithm (O(M2 × N2)) that minimizes the slowest stage in a pipeline, explicitly accounting for master-node overhead (language model head + output layer) and network transmission time between stages.
-
What it can do: Partition 24 transformer blocks across 4 devices in under 1 second, ensuring no single device becomes a bottleneck. In practice, this yields a 2× improvement in decoding throughput (from 8.9 to 16.9 tokens/s) under high demand.
-
Improvement: Use a GA with dual chromosomes—one for device ordering, one for replica grouping—combined with elite preservation (top Q solutions retained) and a 30% mutation rate that includes 4 distinct mutation strategies (swap, regroup, full regroup, full regenerate) to escape local optima.
-
What it can do: Find a deployment plan that processes 22 concurrent decoder requests (vs. 17 for Splitwise) in a 7-device heterogeneous cluster, reducing maximum waiting time from 191s to 75s under 0.5s arrival intervals.
-
Improvement: For each decoder replica, brute-force test batch sizes from 1 to 16 to find the maximum number of parallel requests that still meets a minimum throughput (e.g., 15 tokens/s), then cache this result for reuse.
-
What it can do: Automatically determine that a replica with an Apple M2 Max can handle 10 parallel requests while a Jetson AGX Orin can only handle 2, maximizing overall system throughput without violating Quality-of-Service.
-
Improvement: Analyze the actual input-to-generated token ratio from the incoming request stream (e.g., using a dataset like LongReason) before deployment, and use this ratio to calculate the bottleneck phase (Prefill vs. Decoder) using the formula:
bottleneck = max(NP/PS, ND/DS) - arrival period. -
What it can do: For a code-generation workload with a 0.98 ratio, the system correctly identifies Decoder as the bottleneck and allocates more replicas to it. For a summarization workload with a 2.27 ratio, it shifts resources to Prefill, preventing queue buildup.
-
Improvement: Route each incoming request to the replica that minimizes estimated waiting time, considering both the replica's role (Prefill or Decoder) and its current queue depth, rather than simple round-robin or random assignment.
-
What it can do: Under a 3.0s arrival period, the system achieves a median waiting time of 2.7s (including KV cache transmission) compared to 7.9s for Splitwise, and a P99 of 23.8s vs. 78.0s.
-
Deploy LLMs on edge/fog clusters with mixed consumer hardware (e.g., 7 devices with 6-25GB VRAM each) without cloud infrastructure.
-
Achieve 2× higher decoding throughput (up to 30 tokens/s vs. 12 tokens/s) in low-demand scenarios by exploiting idle capacity.
-
Reduce maximum waiting time by >50% (from 191s to 75s) under high-demand (0.5s arrival) conditions.
-
Adapt to workload changes in real-time by re-optimizing the deployment plan when the input/output token ratio shifts.
-
Guarantee a minimum of 15 tokens/s generation speed per request, which is above the average human reading speed (7-15 words/s), ensuring a natural user experience.
-
Handle 22 concurrent requests in a 7-device cluster, versus 17 for the Splitwise baseline, without additional hardware.
Abstract
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost efficiency, low latency, and optimal resource utilization. Conventional approaches typically assume that an entire model can be hosted on a single device, which does not hold in many real-world scenarios, particularly in Edge and Fog environments where device resources are constrained. In this paper, we introduce E2LLM, a framework designed to enable efficient LLM deployment in such resource limited settings. Rather than simply partitioning a single model across all available devices, E2LLM replicates the full model across multiple groups of devices (replicas) and applies model parallelism within each replica. Each replica is assigned a specialized role PREFILL or DECODER based on its efficiency in handling input and output tokens. This separation leverages the inherent differences between these two phases of LLM inference. To effectively organize devices, we utilize a Genetic Algorithm to form clusters that maximize system performance. Within each cluster, we apply Dynamic Programming to determine an optimal partitioning strategy that minimizes bottlenecks in model-parallel execution. Experimental results demonstrate that our approach adapts robustly to varying workloads, including scenarios with significant variation in input and output token lengths. Compared to the Splitwise baseline, E2LLM reduces average waiting time by over 50% under high-demand conditions
Sources
- Synergistic Tensor and Pipeline Parallelism
- LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
- gpt-oss-120b & gpt-oss-20b Model Card
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing