Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM
summary
The gist
A systematic study of multi-core NPU architectures and LLM serving optimization strategies reveals that careful configuration of hardware, tensor parallelism, and core placement can yield significant
In short
This study systematically explored multi-core NPU architectures for LLM serving by developing a hybrid simulation framework called NpuSim. It analyzed how tensor parallelism, core placement, and memory management affect performance across different hardware setups. The findings guide system design toward specific strategies like heterogeneous PD disaggregation for prefill workloads to achieve significant speedups.
Key concepts
- NpuSim
- A multi-level simulation framework that combines transaction-level modeling with performance models. It simulates memory and interconnect operations at a detailed level while using simpler models for compute operators to balance accuracy and computational speed, allowing it to handle various real-world LLM request patterns.
- Tensor Parallelism Analysis
- The study examined different ways to split the model's computations across multiple cores, such as AllGather and AllReduce. It found that the choice of parallelism (e.g., using AllReduce for K-dimension partitioning) is not always optimal in practice; performance depends heavily on the specific workload.
- Hierarchy Memory Management
- This scheme manages memory for model components like KV cache and weights across different storage tiers: SRAM and HBM. It uses fine-grained management for fast SRAM access (like the KV cache) and coarse-grained management for larger HBM buffers, optimizing allocation based on model size and batch size.
Terminology used across episodes
This episode discusses
- Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM · Paper Radio
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- Qwen Technical Report
- Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies
- WaferLLM: Large Language Model Inference at Wafer Scale
- DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
- MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
- Understanding the planning of LLM agents: A survey
- Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
- DeepSeek-V3 Technical Report
- GPT-4 Technical Report
- Splitwise: Efficient generative LLM inference using phase splitting
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- SCALE-Sim: Systolic CNN Accelerator Simulator
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Gemini: A Family of Highly Capable Multimodal Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
The paper
Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM · Read on arXiv
Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University · Department of Precision Instrument, Tsinghua University
DOI: 10.1109/TCAD.2026.3719164
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM".
Tom: A systematic study of multi-core NPU architectures and LLM serving optimization strategies reveals that careful configuration of hardware, tensor parallelism,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Alright, shifting gears slightly, we need to talk about the title and the authors of this paper, 'Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM'. It’s pretty descriptive, focusing on a systematic study of multi-core architectures and LLM serving optimization strategies.
Jane: The authors are Tianhao Zhu, Dahu Feng, Erhu Feng, and Yubin Xia from the Institute of Parallel and Distributed Systems at Shanghai Jiao Tong University and Tsinghua University Department of Precision Instrument respectively. They seem to be from a strong academic background in parallel systems.
Lu: Their affiliations suggest they have deep roots in the area where hardware architecture meets distributed systems, which makes sense given the focus on multi-core NPUs.
Meng: I’m curious about what their background means for us; are they just academics doing theory, or do they have experience bridging the gap to actual chip design? I need to know if these findings are immediately applicable or if we're looking at a long road.
Lalam: I hope their work helps bridge that gap, because for us, understanding how the underlying hardware operates is essential for building AI that isn't just theoretically sound but practically deployable and fast.
Tom: They’re not just proposing an idea; they are presenting a comprehensive study to address the growing demand for high-performance LLM inference services, which is a huge topic right now.
Jane: The title really emphasizes the "systematic exploration," which tells us this isn't just an anecdotal observation; it’s a thorough investigation into how to optimize these complex AI accelerators.
Lu: That systematic approach is what makes the findings more robust; they aren't relying on a single, perhaps lucky, configuration but are covering a wide area of possibilities.
Meng: If they’ve covered so much ground, I wonder if the practical implications will be finding a simple "magic setting" or if it will require us to build complex automated systems to apply these results.
Lalam: I believe the real value lies in giving us the tools to build that automated system, so we can adapt our AI serving stack as models get bigger and more complex.
The paper's summary: Tom: Now that we’ve talked about the title and authors, let’s look at the actual summary of 'Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM'. Essentially, they explain that they are using a multi-level simulation framework to analyze how hardware configuration, tensor parallelism, core placement, memory management, and the choice between PD disaggregation and PD fusion impacts performance.
Jane: That summary boils down to them using WaferAI-SIM to model memory at the transaction level for high fidelity while using performance models for compute operators to keep things efficient.
Lu: That hybrid simulation approach is key because it balances the need for simulation accuracy with the reality of computational cost, which they achieved by modeling those operators differently.
Meng: It sounds like they’re tackling the problem that many prior works have only looked at one aspect—like just tensor parallelism or core placement—instead, they are looking at the whole system interplay.
Lalam: That holistic view is important because when we're serving a large language model, every single component—from how data moves through the cores to where it lives in memory—is interdependent.
Tom: So, what this means for us is that they’ve mapped out a comprehensive set of variables that we need to consider when optimizing our LLM inference stack, which is a great roadmap.
Jane: It gives us a clear framework to understand where the potential performance gains are hiding, whether it’s in memory allocation ratios or in choosing between different parallelism methods.
Lu: Their analysis of tensor parallelism and core placement strategies shows that the theoretically best strategy often falls short in real-world deployment, which is a very sobering point for us to take away from their findings.
Meng: That suggests we can’t just blindly follow theoretical papers; we need the simulation tools they developed to test those strategies against actual workload distributions.
Lalam: It means our path forward involves leveraging this kind of deep analysis so that our future AI applications are built on optimized foundations, not just hoping for good performance.
The paper's improvements: Tom: Moving into the improvements section, they suggest a few specific things we should focus on: optimizing memory management through a multi-granularity scheme and systematically studying PD disaggregation versus PD fusion strategies for different workload types.
Jane: They propose managing KV cache, weights, and activations across different levels—SRAM at a block granularity and HBM at a coarse granularity—and then determining the optimal allocation ratios based on model size and batch size.
Lu: The idea of fine-grained management for the KV cache in SRAM is very intriguing because it suggests that even small, localized optimizations in memory access can have a big impact on latency.
Meng: That sounds like something we can definitely work on; figuring out how to strategically allocate those buffers between the fast SRAM and the larger HBM based on the model’s specific characteristics is a very concrete engineering task.
Lalam: For me, this memory management directly translates into handling much longer contexts without running into severe memory bottlenecks during inference, which is a major hurdle for complex AI agents.
Tom: And they also give guidance on PD strategies: prioritizing pipeline parallelism in core placement for disaggregation when prefill cores are placed strategically, and using chunked prefill with tensor parallelism for fusion when decoding dominates.
Jane: Their conclusion about which strategy to use based on the workload—heterogeneous disaggregation for prefill-dominated scenarios and PD fusion for decode-dominated workloads—provides a clear rule of thumb we can use in our system design.
Lu: It’s interesting that they also mention specific guidance on tensor parallelism choices depending on the sequence length, which shows that even within a single operation, there are practical trade-offs to consider.
Meng: So, it seems the paper provides actionable advice on how to configure our hardware and software layers to handle different inference scenarios effectively.
Lalam: That actionable guidance is what moves research out of the abstract and into something we can actually implement in production systems, which is incredibly exciting for our culture.
Conclusion: Tom: Alright, we’ve covered a lot about this 'Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM', and now it’s time to wrap up with the main conclusions. They summarize that the overall solution yields performance gains between one point three two times and six point zero three times over state-of-the-art designs, depending on the specific hardware configuration analyzed.
Jane: So, in short, they’ve demonstrated that by carefully configuring hardware, tensor parallelism, and core placement using WaferAI-SIM simulations, you can achieve significant speedups for LLM inference on multi-core NPUs.
Lu: The implication is that we move toward a more predictable way of designing these systems because we have a framework to systematically test configurations rather than just hoping the best design emerges.
Meng: From an engineering perspective, this suggests that our next generation of AI accelerators won't be optimized by chance; they’ll be optimized through rigorous simulation and systematic testing before fabrication.
Lalam: For me, the real impact is seeing a path toward building AI systems that are inherently more efficient, which means lower latency and higher sustained throughput for our users.
Tom: Fantastic points. So to end this discussion on 'Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM', we’ve seen how they systematically explore configurations, propose tailored memory and parallelism strategies, and find performance improvements up to six point zero three times over current SOTA designs.
Jane: It’s a lot of detail, but the core message is that detailed configuration of hardware, tensor parallelism, and core placement makes a real difference in how fast these models run.
Lu: This paper provides a solid foundation for understanding the practical trade-offs involved in deploying large language models on diverse multi-core architectures.
Meng: We’ll take these insights and start feeding them into our design pipeline immediately to see what we can gain practically.
Lalam: I'm just excited for what we can build next; this research paves the way for AI that is not only powerful but also incredibly efficient in deployment.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization