CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM

arXiv:2504.17584 · cs.AR, cs.LG · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM".

Jane: The paper was written by Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du et al. from Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper with a pretty dense title: “CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM.”

Jane: And Tom, I have to say, that title packs a lot in. Let’s unpack it for our listeners. “Attention-FC Disaggregated” means we’re splitting the brain of an AI model into two parts—the part that figures out what to focus on, and the part that does the heavy math.

Tom: Right, and the paper argues that these two parts have totally different needs. One is starving for memory space, the other is starving for compute speed. So why keep them stuck together on the same chip?

Jane: Exactly. And that’s where “DIMM-PIM” comes in. That’s a fancy way of saying we’re putting little computers right inside the memory sticks of a server, the same kind of sticks you’d find in a regular PC, just way bigger.

Tom: So instead of shipping data back and forth to the GPU, you do the “attention” work right where the data lives. It’s like reading a book in the library instead of checking it out and carrying it home.

Jane: I love that analogy. And the team behind this is from Shanghai Jiao Tong University. They’ve got a really bold claim here—they say their system can be over five times faster than the previous best approach.

Tom: Five times faster for these long-context models, which are the ones that can read entire novels or write long code files. That’s a huge deal.

Jane: It really is. And what’s clever is they didn’t just invent new hardware. They built a whole system around it, which is why the paper is so substantial.

Tom: So, Jane, you’re telling me this isn’t just a cool lab experiment? This could actually change how companies run their AI?

Jane: That’s the promise. And I’m really curious to see how they pulled it off, because making memory sticks compute things is notoriously tricky.

Tom: Well, stick around, because we’re going to break down exactly how they did it, starting with their big-picture model of the problem.

Summary: Tom: So we’ve got the title sorted. Now, Jane, what’s the one-sentence summary of this paper that we can hold in our heads?

Jane: I’d say it’s this: they built a system that makes long-context AI inference both faster and cheaper by putting the right kind of compute in the right kind of memory.

Tom: And that’s a big deal because, as the paper points out, we’re hitting a wall. When you ask an AI to handle a huge amount of context, the memory becomes the bottleneck, not the processor.

Jane: Right. They call it the “Liebig’s Law” of AI systems. It’s an old agricultural idea—a plant grows only as fast as the scarcest nutrient allows. Here, the “nutrients” are memory capacity and memory bandwidth.

Tom: So if you have tons of memory space but slow access, you’re stuck. And if you have fast access but not enough space, you’re also stuck. You have to balance both.

Jane: Exactly. And that’s the core insight. Previous attempts to fix this problem only focused on one side. They either added more memory or made the memory faster, but never both at the same time.

Tom: So their solution, CHIME, is specifically designed to scale both. They’re using those memory sticks with built-in processors we talked about earlier.

Jane: And the results speak for themselves. They’re seeing up to a five point one five times speedup over the previous state-of-the-art systems that used a different type of memory called HBM.

Tom: That’s not a small improvement. That’s a game-changer for anyone running these massive models.

Jane: It is. But the paper isn’t just about the hardware. They also had to write clever software to manage the communication between the GPU and these new memory sticks.

Tom: Because otherwise, you’d spend all your time waiting for data to move around, right?

Jane: Precisely. They have a whole scheduling system to keep everything busy and avoid those idle gaps. It’s a full-stack solution.

Tom: So it’s not just a chip, it’s a whole way of thinking about the problem.

Jane: Exactly. And that’s what makes it so interesting. They’re not just tweaking a design; they’re proposing a new paradigm for how to build AI inference systems.

Tom: I’m excited to get into the nitty-gritty of how they actually made this work.

Improvements: Tom: Okay, so we know CHIME is fast. But what are the actual improvements they’re proposing over what came before?

Jane: The biggest one is the hardware itself. They’re using DIMM-PIM, which is a way of putting processing power directly onto standard memory modules.

Tom: And that’s different from the previous approach, which was to use HBM-PIM, right?

Jane: Right. HBM is that super-fast, super-expensive memory that sits right next to the GPU. It’s great for speed, but it’s limited in capacity and costs a fortune per gigabyte.

Tom: So the trade-off is speed versus space.

Jane: Exactly. And for long-context models, you need space. The model has to remember everything you’ve said, and that takes up a lot of memory. HBM just can’t hold it all.

Tom: So they moved to DIMM, which is the standard, much larger memory. But standard DIMM is slow.

Jane: That’s the trick. They’re taking that standard, large memory and adding processing units right next to the memory banks. This gives them the capacity of DIMM with a massive boost in bandwidth.

Tom: So they get the best of both worlds. But I’m guessing it’s not as simple as just gluing a processor to a memory stick.

Jane: You’d be right. The paper spends a lot of time on the challenges. For example, when you have multiple memory chips working together, you have to synchronize them. If they’re not perfectly in sync, you get bubbles—idle time where nothing is happening.

Tom: Bubbles are bad.

Jane: Very bad. So they designed a “bubble-free” pipeline to keep the data flowing smoothly. They also had to solve a data layout problem, making sure the data is arranged in the memory in a way the processors can actually use.

Tom: So it’s a hardware and software problem.

Jane: It’s a co-design problem. The hardware is designed with the software in mind, and the software is written to get the most out of the hardware.

Tom: And they also improved the scheduling on the software side, right?

Jane: Yes. They have a scheduler that predicts how long operations will take on the GPU versus the memory sticks, and it balances the workload to keep both busy. It’s like a traffic controller for compute tasks.

Tom: So instead of one device waiting for the other, they’re working in parallel.

Jane: Exactly. That’s how they squeeze out that extra performance. It’s a really holistic approach.

Tom: I’m starting to see why this paper is getting so much attention.

First Page: Tom: Let’s zoom in on the very first page of “CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM.” It starts with a quote about Liebig’s Law.

Jane: That’s the plant growth analogy we mentioned. And they use it to set up their core argument: you can’t just improve one thing and expect the whole system to get faster.

Tom: They even show a concrete example. They simulated a huge model called GPT-175B and found that making the memory bandwidth sixteen times faster only improved the overall speed by less than one percent.

Jane: That’s a stunning result. It really proves their point. All that extra bandwidth was useless because the system was limited by something else—in that case, memory capacity.

Tom: So they’re saying you have to look at the whole system, not just the individual parts.

Jane: Exactly. And that’s why they built their own model, the Disaggregated Roofline Model, to analyze the whole system. It helps them figure out where the real bottleneck is.

Tom: And based on that analysis, they decided that DIMM-PIM was the right choice because it offers a more balanced configuration of capacity and bandwidth.

Jane: Right. They also point out the economic angle. HBM memory is over six times more expensive per gigabyte than the standard DIMM memory they’re using.

Tom: So not only is it faster in the scenarios that matter, but it’s also cheaper to build.

Jane: That’s the dream. Better performance and lower cost. And that’s what makes this paper so compelling.

Tom: They also mention that the DIMM interface is more scalable. You can just plug in more memory sticks to get more capacity.

Jane: It’s a more flexible and future-proof design. As models get bigger and contexts get longer, you can just add more hardware without redesigning the whole system.

Tom: So the first page really sets the stage for a very practical, well-reasoned approach to a huge problem.

Jane: It does. And it’s a great example of how a simple analogy from agriculture can lead to a breakthrough in computer architecture.

Tom: I love when that happens. Science is all about connecting ideas.

Conclusion: Tom: We’ve covered a lot of ground on “CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM.” Let’s wrap it up.

Jane: The big takeaway is that they’ve identified a fundamental law—you need to balance memory capacity and bandwidth to get the best performance out of AI inference.

Tom: And they built a complete system, from the hardware to the software, that actually follows that law.

Jane: Their hardware, CHIME-PIM, puts processing power in standard memory sticks, giving you the space you need without sacrificing speed.

Tom: And their software, CHIME-sys, keeps everything running smoothly, hiding the communication delays and balancing the workload.

Jane: The result is a system that’s not only faster—up to five point one five times faster than the previous best—but also more cost-effective.

Tom: It’s a really elegant solution to a problem that’s only going to get more important as AI models continue to grow.

Jane: Absolutely. And it shows that sometimes the best way forward isn’t to build a faster, more expensive chip, but to be smarter about how we use the memory we already have.

Tom: Well said, Jane. This is definitely a paper that’s going to influence how people design AI infrastructure in the coming years.

Jane: I agree. It’s a new way of thinking about the problem, and it’s a great example of hardware-software co-design done right.

Tom: Alright, that’s a wrap on CHIME. Thanks for joining us, and we’ll see you next time for another deep dive into the world of AI research.

Jane: See you then, everyone!

Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen

Shanghai Jiao Tong University

cs.AR, cs.LG

Submitted: 2026-08-07

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: The paper introduces CHIME, the first Attention-FC Disaggregated (AFD) LLM inference system integrating DIMM-PIM.

Key concepts

Attention-FC Disaggregated Inference
This means splitting an AI model into two parts: the attention mechanism that figures out what to focus on, and the part that performs the heavy mathematical calculations. The paper argues these two parts have different needs for memory space and compute speed.
DIMM-PIM
This refers to putting little computers directly inside standard memory sticks in a server. This allows the attention work to happen right where the data is stored, rather than moving data back and forth to a separate GPU, improving efficiency.
Liebig’s Law of AI Systems
This agricultural idea states that plant growth is limited by the scarcest nutrient. In AI systems, this means performance is limited by either memory capacity or memory bandwidth; you must balance both for optimal results.

Terminology

Summary

The paper introduces CHIME, the first Attention-FC Disaggregated (AFD) LLM inference system integrating DIMM-PIM. AFD systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. The authors first design a Disaggregated Roofline Model (DRM) to characterize AFD performance, revealing that system throughput is constrained by the accelerator's limiting factor: either memory bandwidth or capacity. They observe that prior AFD systems often overlook these constraints and fail to balance them, leading to resource underutilization or constrained throughput. CHIME addresses synchronization challenges inherent to distributed cooperating DRAM chips in DIMM-PIM through bubble-free pipelining and hybrid-grained re-layout for efficient attention computation, and maximizes cross-device resource utilization via rankset-granular communication-computation overlapping and alignment-predicting scheduling. Evaluations show CHIME achieves up to 5.15× speedup over state-of-the-art HBM-PIM solutions.

The paper notes that the demand for Large Language Models (LLMs) to process and generate longer sequences is rapidly increasing, driven by applications like long reasoning and complex code generation. In long-context scenarios, the decoding process becomes a significant performance bottleneck, and the increasingly longer KV cache imposes growing memory capacity and bandwidth demands on the inference system.

AFD systems "partition the LLM inference workload: the memory-intensive Attention operations and KV cache are offloaded to a specialized memory-rich device (hereafter referred to as the 'accelerator'), while the compute-intensive Fully-Connected (FC) operations remain on the GPU or NPU. The paper reveals a critical yet counter-intuitive phenomenon: simply enhancing the accelerator does not always guarantee better system performance. In an evaluation on GPT-175B with OpenR1, equipping DGX-A100's HBMs with PIM from 1× to 16× higher bandwidth only leads to <1% throughput improvement."

The paper designs DRM as the first general performance model to analyze how AFD systems perform in various scenarios. DRM holistically characterizes the performance of both the accelerator and the GPU, crucially modeling their interdependencies. The model concludes a principle: "the system's throughput is limited by the scarcest resource, be it the accelerator's memory bandwidth or capacity. Crucially, improvements to any non-bottleneck resource will result in diminishing or even negligible returns."

Key implications derived from DRM include:

  • Implication I: The overall inference throughput of AFD systems is determined by the lower throughputs of the two devices that separately execute the FC and attention.

  • Implication II: The bandwidth of the accelerator limits the throughput of AFD systems when the attention becomes the bottleneck.

  • Implication III: The capacity that stores the KV cache limits the overall throughput of AFD systems when the FC becomes the bottleneck.

The paper formulates the Liebig's Law for AFD systems: The throughput of an AFD system is constrained by the weaker point in terms of either memory capacity or bandwidth.

The paper argues that PIM integrated with DIMM, i.e., equipping host memory modules with bank-level processing units, constitutes a case of a new accelerator that strikes this balance. Compared to prior HBM/GDDR-PIM solutions, DIMM-PIM offers three key advantages:

  1. It directly alleviates the capacity bottleneck of HBM-PIM while providing significantly higher bandwidth than standard host memory, enabling higher overall throughput.

  2. It inherits the superior scalability and configurability of the standard DIMM interface, making the system adaptable and future-proof for diverse memory demands.

  3. It offers higher economic efficiency, which could be cheaper than HBM/GDDR-PIM and suffer less memory wastage.

CHIME-PIM integrates co-operated processing units (PUs) across both bank and rank levels. The bank PUs are integrated near DRAM banks with DRAM process, fetching KV cache from banks and performing score and context computation, while a shared buffer is leveraged to broadcast input vectors to all bank PUs. The Softmax unit, adder unit, and re-layout unit are integrated on the buffer chip with logic process as a part of the rank PU.

To eliminate data-synchronization overhead, CHIME-PIM leverages decoupled memory buses (internal buses servicing bank PUs and external buses handling rank PU access) and FlashAttention's kernel fusion technique to orchestrate fine-grained kernel pipelining to maximize concurrency while constraining the intermediate head footprint on rank PUs. The paper provides quantitative analysis showing that the communication time is:

Tcomm = (Lt × Ngqa × Nchips)/Brk

And the computation time is:

Tcomp = (Lt × Eh × ⌈Ngqa/Ncmr⌉)/(Bbk × Nbk × Nhc)

The condition for bubble-free execution is Tcomm ≤ Tcomp, which requires head mapping satisfying:

Nhc ≤ (Eh × Brk × ⌈Ngqa/Ncmr⌉)/(Bbk × Nbk × Ngqa × Nchips)

The paper applies Nhc = 8 for MHA and Nhc = 1 for GQA-8.

The paper addresses layout mismatch through hybrid-grained relayout, which leverages the rank PU's re-layout unit to perform in-flight data transformation during communication. It performs fine-grained re-layout ensuring each element resides on a single chip and coarse-grained re-layout mapping each head to Nhc chips.

The paper identifies that due to the shared memory buses, only one rank can be accessed in a channel at the same time during communication, while other ranks remain idle. It proposes the rankset, which is composed of one rank from every channel, forming the basic granularity of independent communication and computation. During communication of one rankset, other ranksets can perform independent computation without blocking, which could preserve 2/3 of computational power during communication.

CHIME's scheduler models and predicts the execution latencies on the two devices that helps to align the parallel execution latencies. It selects requests to form sub-batches, whose predicted latencies on the two devices are aligned. The latencies are modeled as:

TGPU0 = tp(cp0, fp0) + tbatch(cp0, fd1)

TPIM0 = td(fd0) + tcomm(fd0, cp1)

TGPU1 = tp(cp1, fp1) + tbatch(cp1, fd0)

TPIM1 = td(fd1) + tcomm(fd1, cp0)

The scheduling policy: (1) adds one prefilling request into each sub-batch; (2) adds N decoding requests to each sub-batch, predicting TPIM and TGPU, continuing until TPIM > TGPU; (3) repeats until PIM memory is exhausted, with the last prefilling request chunked to align TGPU with TPIM.

For modeling, CHIME uses Random Forest Regression (RFR) for GPU latency prediction and a simple and fast linear model for PIM execution, since CHIME-PIM execution is featured predictable (i.e., execution time is linearly related to the number of computed/transferred tokens).

CHIME achieves up to 5.15× higher throughput than the HBM-PIM baseline, 3.45× than the HBM-PIM-EXT baseline, 3.94× than the GPU-only baseline, and 7.21× higher than R-PIM baseline. The improvements are attributed to much larger batch sizes compared to the HBM-based baselines (CHIME has 2TB of host memory for KV cache storage versus about 310GB for GPU/HBM-PIM baselines).

Compared with GPU and HBM-PIM, CHIME increases the batch size by 6.6×, while the latency per batch increases by only 2.2×, indicating the throughput improvement of CHIME primarily stems from the increased batch size and the corresponding improvement in GPU utilization.

"Expanding either bandwidth or capacity alone does not effectively improve throughput. For example, when exclusively scaling memory capacity or bandwidth by 8× on GPT-175B, the throughput only increases by 2.28× and 1.01×, respectively. In contrast, when both bandwidth and capacity are enlarged, the throughput increases by 8.23×."

  • Bubble-free pipelining achieves about 27.9% and 74.4% latency reduction on MHA and GQA computation, respectively.

  • Hybrid re-layout enables up to 17% latency reduction.

  • Scheduler can significantly reduce latency by up to 70.93% without sacrificing the throughput.

  • Rankset-granular overlapping could reduce the overhead by up to 75.08%.

CHIME achieves a 40% total energy reduction, primarily by executing FC on the GPU with larger batch sizes, which reduces the number of weight loading. The paper also notes that the per-GB cost of HBM2e can be over 6× higher than DDR4.

The paper presents "the first general AFD performance model, Disaggregated Roofline Model, from which we conclude the 'Liebig's Law' that guides the design of CHIME, a hardware-software co-designed AFD system with DIMM-PIM that offers scalable memory capacity and bandwidth for LLM inference. CHIME can enable efficient DIMM-PIM attention computation and maximizes the resource utilization for cross-device inference. Evaluations show CHIME achieves significantly higher throughput, establishing a new paradigm for efficient long-context LLM inference."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with the resulting capabilities:

Improvement: I will build a performance analysis tool that models the throughput of Attention-FC Disaggregated (AFD) inference systems by jointly considering GPU FC throughput and accelerator attention throughput, including memory bandwidth and capacity constraints.

Resulting Capability: The AI system can predict, before deployment, whether an accelerator (e.g., HBM-PIM, DIMM-PIM, CPU) will be the bottleneck and whether scaling bandwidth or capacity alone will yield diminishing returns. This prevents costly over-provisioning of resources (e.g., buying excessive HBM bandwidth when capacity is the limit) and guides hardware selection for specific workloads (e.g., long-context vs. short-context).

Improvement: I will design the system to offload memory-bound attention operations to DIMM-PIM (bank-level processing units on host memory) instead of HBM-PIM or CPU, leveraging its scalable capacity (2TB) and high bandwidth (13 TB/s) to balance the memory bottleneck.

Improvement: I will implement a pipelined execution of attention (score, softmax, context) that overlaps cross-chip data transfer with bank-PU computation, using quantitative analysis to determine optimal head-to-chip mapping (e.g., Nhc=8 for MHA, Nhc=1 for GQA-8).

Improvement: I will implement a re-layout unit that performs both fine-grained (bit-level) and coarse-grained (element-level) data transformation during QKV offloading/onloading, ensuring each element resides on a single DRAM chip and heads are mapped to the correct chips.

Improvement: I will schedule data transfers (QKV offload, output onload) at the granularity of ranksets (one rank per channel), allowing other ranksets to compute attention concurrently, and use layer-interleaved KV cache storage for load balance.

Improvement: I will use a scheduler that models execution latencies on both GPU (via Random Forest Regression) and CHIME-PIM (via linear models) to select requests forming sub-batches whose latencies are aligned, preventing idle bubbles.

Improvement: I will design the system to scale memory capacity and bandwidth simultaneously (e.g., adding more DIMM-PIM ranksets) rather than scaling one dimension alone, as per Liebig's Law.

Improvement: I will integrate runtime profiling and incremental learning (RFR) to predict GPU and PIM latencies with <1% relative error, enabling dynamic adaptation to changing workloads.


Overall, the improved AI system can: Serve long-context LLM inference (e.g., 12K+ output tokens) with up to 5.15× higher throughput than state-of-the-art HBM-PIM, support batch sizes 6.6× larger, reduce energy by 40%, and scale efficiently by balancing memory capacity and bandwidth—all while maintaining predictable, low-latency execution.

Abstract

Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. In this paper, we first design a Disaggregated Roofline Model (DRM) to characterize AFD performance, revealing that system throughput is constrained by the accelerator's limiting factor: either memory bandwidth or capacity. We observe that prior AFD systems often overlook these constraints and fail to balance them, leading to resource underutilization or constrained throughput. Therefore, we propose CHIME, the first AFD system integrating DIMM-PIM, which is a case of the new accelerator that strikes the balance with scalable capacity and bandwidth. To address the synchronization challenges inherent to the distributed cooperating DRAM chips in DIMM-PIM, CHIME employs bubble-free pipelining and hybrid-grained re-layout for efficient attention computation. Furthermore, it maximizes cross-device resource utilization via rankset-granular communication-computation overlapping and alignment-predicting scheduling. Evaluations show CHIME achieves up to 5.15 times speedup over state-of-the-art HBM-PIM solutions.

Sources

Related papers