WWW: What, When, Where to Compute-in-Memory
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "WWW: What, When, Where to Compute-in-Memory".
Jane: The paper was written by Tanvi Sharma, Mustafa Ali, Indranil Chakraborty and Kaushik Roy from Purdue University and Microsoft and Google.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone! Today we're diving into a paper that's been making the rounds on arXiv, and it's got one of those titles that just makes you stop and think: "What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication during Machine Learning Inference." Jane, I gotta say, just reading that title out loud feels like the paper is asking the universe three very specific questions.
Jane: Tom, it really does! And I love that framing, because for years we've been hearing about compute-in-memory as this magic bullet for making AI faster and more efficient. But this paper is essentially saying, "Hold on, not all compute-in-memory is created equal, and not every situation calls for it." It's like asking whether you should use a sports car or a pickup truck—it depends on what you're hauling.
Tom: Exactly! And the authors—Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, and Kaushik Roy from Purdue—they're not just theorizing. They've built an analytical framework to actually answer those three questions. The "what" is about the type of compute-in-memory primitive, the "when" is about the workload shape, and the "where" is about which level of the memory hierarchy you integrate it into.
Jane: Right, and that "where" part is so crucial. We're talking about a GPU-like architecture, with register files and shared memory. The paper is essentially saying, "If you're going to put compute inside memory, you need to know if it belongs in the tiny, super-fast register file or the bigger, slightly slower shared memory." And the answer isn't obvious until you actually do the math.
Tom: And the math is pretty compelling. They show energy efficiency improvements up to three point four times and throughput improvements up to fifteen point six times compared to a baseline tensor-core architecture. But here's the kicker—those numbers aren't universal. They depend heavily on the shape of the matrix multiplication you're doing.
Jane: That's the "when" question. Some matrix multiplications are like a big, square block—lots of reuse, lots of compute. Others are long and skinny, like a matrix-vector multiply, and those don't benefit nearly as much. In fact, for some of those skinny shapes, the compute-in-memory approach can actually be worse than just using regular cores.
Tom: So the paper isn't just cheerleading for compute-in-memory. It's giving us a roadmap for when it's actually worth the integration effort. And that's the kind of nuanced, practical insight that I think the hardware community really needs right now.
Jane: Absolutely. And it makes me wonder—what does this mean for the next generation of AI accelerators? Are we going to see chips that dynamically decide whether to use compute-in-memory or standard cores depending on the layer of the neural network?
Tom: That's a great question, and I think the paper's framework could absolutely inform that kind of adaptive design. But let's not get ahead of ourselves. We need to talk about the actual methodology they used to get these results. That's coming up next.
Summary: Tom: So Jane, we've set the stage with those three big questions—what, when, where. Now let's talk about how this paper actually goes about answering them. The core idea is that they've built a dataflow-centric way of representing compute-in-memory primitives.
Jane: Right, and I think that's a really clever move. Instead of getting bogged down in the transistor-level details of every single SRAM design out there, they abstract it into something the architecture can understand. They break a compute-in-memory primitive into smaller "CiM units," and each unit has a certain number of rows and columns it can process in parallel, and a certain number it processes sequentially.
Tom: Exactly. So you have these parameters—Rp and Cp for parallel rows and columns, and Rh and Ch for the sequential hold factors. That lets them compare very different designs, like an analog 6T SRAM primitive versus a digital 8T one, all under the same analytical umbrella.
Jane: And that's important because these primitives are wildly different. The analog ones are super energy-efficient per operation, like zero point zero nine picojoules for an eight-bit MAC, but they're slow—one hundred forty-four nanoseconds. The digital ones are faster but use more energy. Without a unified framework, you'd be comparing apples to oranges.
Tom: Right, and they also enforce iso-area constraints. So if one primitive takes up more silicon area, you get fewer of them in the same cache space. That's the fair way to compare—you're asking, "Given the same chip area, which design gives me the best performance?"
Jane: And then there's the mapping algorithm. This is where they really shine. They've written a priority-based algorithm that decides how to map a given GEMM—that's general matrix multiplication—onto the compute-in-memory primitives. The first priority is keeping the weights stationary, because that's the whole point of compute-in-memory. You want to load the weights once and then stream the inputs through them.
Tom: And they compare this to a heuristic search approach, which is what a lot of other tools use. Their algorithm is not only faster to run, but it consistently finds better mappings—up to six point six times better hardware utilization, and about one point two times better energy efficiency.
Jane: That's a big deal. It means their approach isn't just a theoretical exercise. It's a practical tool that could be used in real chip design flows to quickly evaluate whether a particular compute-in-memory design is worth pursuing.
Tom: And speaking of practical, the results they get are really interesting. They test on real workloads—ResNet50, BERT-Large, GPT-J, DLRM—and they find that the benefits vary a lot. BERT-Large, with its big, regular matrix multiplications, gets huge gains. But some of the GPT-J layers, especially during the decoding phase, have these skinny matrix-vector multiplies, and those don't benefit at all.
Jane: Right, and that brings us back to the "when" question. The paper is essentially saying, "Don't just bolt compute-in-memory onto everything. Look at your workload first." And that's such a practical, engineering-minded approach.
Tom: It really is. So we've got the framework, we've got the mapping algorithm, and we've got the workload analysis. But the real meat is in the "where" question—where in the memory hierarchy should you put this thing? That's what we're going to dig into next.
Improvements: Tom: Alright, Jane, let's get into the "where" question, because this is where the paper really earns its keep. They compared integrating compute-in-memory at two levels: the register file, which is tiny and super fast, and the shared memory, which is bigger and slower but has way more capacity.
Jane: And the results are fascinating. At the register file level, you get decent improvements—about three times better energy efficiency for BERT layers compared to a baseline tensor core. But the throughput is capped because you only have so many compute-in-memory primitives fitting in that small space.
Tom: Right, and then they look at shared memory. They actually consider two configurations. One keeps the same number of compute primitives as the register file version, just to isolate the effect of the memory level itself. The other fills the entire shared memory with compute primitives, which is a much bigger setup.
Jane: And that second configuration is where things get wild. The throughput jumps to about ten times what the register file version achieves. Because you've got so many more primitives working in parallel. But here's the twist—the energy efficiency doesn't improve as much as you'd hope.
Tom: Yeah, that surprised me too. The energy efficiency at the shared memory level is actually lower than at the register file level when you keep the same number of primitives. Because without that intermediate register file to cache data, you're going to main memory more often, and that's expensive.
Jane: But when you fill the whole shared memory with compute primitives, the energy efficiency does go up—about zero point two five TOPS/W better than the register file version. So the capacity of the memory level matters more than its position in the hierarchy.
Tom: That's a really important takeaway. It suggests that if you're designing a chip for compute-in-memory, you should focus on maximizing the amount of on-chip memory that can do computation, rather than trying to put compute in every single level.
Jane: And they also compare against the baseline tensor core architecture. The compute-in-memory at shared memory level gets up to fifteen point six times better throughput for some workloads. But for those skinny matrix-vector multiplies, it's actually worse. The baseline tensor core can be more flexible with its dataflow, while compute-in-memory is locked into a weight-stationary approach.
Tom: So the paper is really saying, "Here's the sweet spot, and here's where you should avoid." That's the kind of guidance that chip designers desperately need, because compute-in-memory is a significant investment in terms of design effort and area.
Jane: Absolutely. And it also points to future work—like, could you have a hybrid architecture that dynamically switches between compute-in-memory and standard cores depending on the layer? That would be the ultimate answer to the "when" question.
Tom: That's a fantastic vision. And it's not just academic—the paper's framework could be used to evaluate exactly that kind of hybrid design. But before we get too far into the future, let's wrap up what we've learned today.
Conclusion: Tom: Well, Jane, we've covered a lot of ground on "What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication during Machine Learning Inference." Let's pull it all together.
Jane: Absolutely. The paper gives us three clear answers. For "what," the digital 6T SRAM primitive with an adder tree design offers the best balance of throughput and energy efficiency. For "when," compute-in-memory shines with large, regular matrix multiplications—like those in BERT—but struggles with skinny matrix-vector multiplies. And for "where," the shared memory level, when fully packed with compute primitives, delivers the biggest performance gains.
Tom: And the key insight that ties it all together is that capacity matters more than hierarchy position. The more compute you can pack into on-chip memory, the better, because you're reducing those expensive trips to main memory.
Jane: Right. And they've also given us a practical mapping algorithm that's faster and more effective than existing heuristic searches. That's a tool that could be used in real chip design flows today.
Tom: The paper also highlights a sobering reality—compute-in-memory isn't a silver bullet. For memory-bound workloads with low reuse, it can actually underperform a standard tensor core. So the design guidance is nuanced, which is exactly what engineers need.
Jane: And the implications are huge. As AI models get bigger and more complex, the energy cost of moving data around becomes the dominant bottleneck. This paper gives us a roadmap for mitigating that, at least for the matrix multiplication operations that form the backbone of inference.
Tom: So, to the authors—Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, and Kaushik Roy—great work. You've given the hardware community a clear, actionable framework for integrating compute-in-memory where it actually helps.
Jane: And to our listeners, if you're designing AI hardware, this paper should be on your reading list. It's not just theory; it's a practical guide with real numbers and real insights.
Tom: Alright, that's a wrap on "What, When, Where to Compute-in-Memory." Next up, we've got a paper on spiking neural networks that's been getting a lot of buzz. Until then, keep computing, and maybe do it in memory—but only when it makes sense!
Jane: See you next time, everyone!
Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, Kaushik Roy
Purdue University · Microsoft · Google
cs.AR, cs.DC, cs.LG
Submitted: 2025-02-27
Updated: 2026-08-18
Comments: added supplementary
Code: https://github.com/tanvisharma/WWWtoCiM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 76/100
The gist: This paper addresses the challenges of integrating Compute-in-Memory (CiM) paradigms into on-chip memory subsystems for efficient matrix multiplication during Machine Learning (ML) inference.
Key concepts
- Compute-in-Memory (CiM)
- This refers to placing computation directly inside the memory units, such as SRAM. It aims to make AI faster and more efficient by reducing data movement between memory and processors. The paper analyzes different CiM primitives like analog versus digital designs.
- Workload Shape
- This refers to the structure of the matrix multiplication being performed, such as whether it is a large, square block or a long, skinny matrix-vector multiply. The benefits of compute-in-memory depend heavily on this shape; some shapes benefit greatly while others do not.
- Memory Hierarchy Levels
- This refers to the different storage tiers in a computer system, specifically comparing the tiny register file and the larger shared memory. The paper investigates where to place CiM units within these levels to maximize performance gains.
- Mapping Algorithm
- This is a priority-based algorithm developed by the authors that decides how to map general matrix multiplications onto CiM primitives. It prioritizes keeping weights stationary, and it was found to be faster and more effective than existing heuristic search approaches.
Terminology
Summary
This paper addresses the challenges of integrating Compute-in-Memory (CiM) paradigms into on-chip memory subsystems for efficient matrix multiplication during Machine Learning (ML) inference. The authors systematically investigate three fundamental questions: What type of CiM to use,
When to use CiM,
and Where to integrate CiM
within the memory hierarchy.
Motivation and Problem Statement: The paper highlights that matrix multiplication (GEMMs) is the dominant computation during ML inference. Due to the separation of processing units from memory in von Neumann architectures, such data-intensive multiplications incur high energy costs from frequent memory accesses, known as the memory wall
or von-Neumann bottleneck.
While CiM paradigms have emerged as energy-efficient solutions, integrating them poses key questions: 1) What type of CiM to use given a multitude of design characteristics, 2) When to use CiM given workloads with varied memory and compute requirements, and 3) Where to integrate CiM given that each memory level has different bandwidth and capacity.
Approach: The authors develop a systematic methodology for fair comparison across various CiM primitives, workloads, and memory levels. They devise a dataflow-centric representation of CiM primitives by breaking down designs into smaller CiM units, abstracting CiM specifications in terms of compute parallelism and compute energy. They propose a priority-based dataflow mapping algorithm that aims to maximize weight-stationary benefits of CiM and reduce data movement. The evaluation assumes iso-area constraints and considers typical CiM primitives at different cache levels for various GEMM shapes based on real and synthetic workloads.
Key Contributions:
-
Analytical evaluation of SRAM-based CiM primitives at Register File (RF) and Shared Memory (SMEM) levels in a tensor-core-like processor architecture.
-
A dataflow algorithm to optimize performance and energy efficiency through optimal mapping for a given CiM architecture and GEMM shape.
-
Detailed analyses answering what, when, and where to use CiM for various GEMM shapes from energy/performance perspectives.
CiM Primitive Representation: The paper introduces a dataflow-centric approach where a CiM primitive is divided into Rp × Cp CiM units, all operating in parallel. Each CiM unit can sequentially perform Rh × Ch MAC operations. The parameters Rp/Cp (row/column parallelism) and Rh/Ch (row/column hold factors) capture inherent limitations in row and column parallelism within a CiM primitive. The approach can be extended to different types and numbers of CiM primitives.
Mapping Algorithm: The priority-based mapping algorithm:
-
First priority: Keep weights stationary to reduce writes, mapping reduction dimension (K) to rows and output columns (N) to columns of all CiM primitives.
-
Second priority: Maximize input/weight reuse by mapping the maximum possible input matrix (M × K) to adjacent memory levels.
-
For multiple CiM primitives, priority is given to higher parallelism (using multiple primitives simultaneously) over fully utilizing each CiM unit.
-
Loop order is decided to leverage inherent temporal reductions, with M in the innermost loop for highest input reuse, and K faster than N to prioritize reducing output partial sums.
The algorithm consistently outperforms heuristic mapping search, achieving more than 6.6× improvements in utilization on average, 1.2× improvement in energy-efficiency, and more than 3.2× enhancement in throughput across different GEMM types.
Experimental Setup: The baseline architecture consists of a single SM connected to main memory, with each sub-core containing 16×16 processing elements (PEs) representing tensor-core-like operations. INT-8 precision, 45nm technology, and 1GHz frequency are assumed. The cache capacities are 4×4 KB for RF and 256 KB for SMEM. Four CiM primitives are evaluated: Analog-6T, Analog-8T, Digital-6T, and Digital-8T, with varying specifications for energy, latency, parallelism, and area overhead.
Key Results:
Effect of CiM Primitive Choice: The CiM primitive with lowest energy cost (Analog-8T at 0.09pJ) shows highest energy efficiency at system level, achieving more than 3 TOPS/W. However, the advantage in energy efficiency is significantly lower than standalone CiM accelerators due to expensive main memory accesses. The Digital-6T primitive with adder tree design achieves the highest throughput due to its ability to exploit full row and column parallelism. Analog primitives suffer from lower performance due to row/column multiplexing techniques.
Effect of Workload Choice:
-
For weight matrix (N=K): Energy efficiency increases until dimension K reaches on-chip memory capacity (capped at 512), then plateaus at approximately 1.75 TOPS/W. For fixed weight size, energy efficiency increases with M up to a sweet point, then falls due to increased main memory accesses.
-
For input matrix (M=K): Energy efficiency initially increases with matrix dimensions but tapers off after M=K=512. Increasing N positively affects energy efficiency regardless of input matrix size.
-
For output matrix (M=N): TOPS/W rises with output matrix dimensions but with diminishing returns. K dimension shows an optimal point at K=256 (maximum rows Digital-6T can process simultaneously), beyond which TOPS/W declines due to increased partial sum accesses.
Effect of Memory Level:
-
Register File integration: BERT-Large layers achieve exceptional energy efficiency (>1.67 TOPS/W) and highest throughput (455 GFLOPS) due to large, regular GEMM shapes. Matrix-vector multiplications (M=1) suffer from low energy efficiency (0.03 TOPS/W) and throughput (≈31 GFLOPS).
-
Shared Memory integration: ConfigB (using all CiM primitives that fit in SMEM under iso-area constraints) significantly enhances throughput, exceeding RF performance by approximately tenfold, and achieves higher energy efficiency by ≈0.25 TOPS/W compared to RF.
Comparison with Baseline: On average, CiM integration at RF improves energy efficiency by up to 3.4× and throughput by up to 15.6× compared to the baseline tensor-core architecture with INT-8 precision. BERT layers consistently derive the highest benefits (3× increase in TOPS/W). However, CiM integration is disadvantaged when M dimension is extremely small, where the baseline's flexible dataflow allows better hardware utilization.
Key Takeaways:
-
Maximum throughput gain is achieved by Digital-6T for medium to large GEMM shapes.
-
Analog-8T achieves maximum energy reduction under iso-area constraints.
-
CiM integrated caches do not increase performance of memory-bound layers.
-
GEMMs with high K value have high energy and performance benefits from CiM integration; small K GEMMs achieve better throughput with baseline.
-
Highest performance gains are observed at SMEM level compared to RF level under iso-area constraints.
-
The storage capacity of memory is more important than its hierarchical level when integrating CiM.
Recommendations: The authors recommend that CiM offers the best solution for ML inference in terms of energy-efficiency but lags in performance due to limited parallelism and high latency. Designing larger CiM accelerators that can map all matrix dimensions onto on-chip memory is more advantageous. Memory level with highest capacity is the optimal level to integrate CiM. The size of CiM primitive-based accelerators should be tailored to accommodate workload-specific reductions in dimension K. Avoid deploying CiM primitives for matrix-vector multiplication tasks where data reuse is minimal.
Improvements for AI systems
Based on the paper, here are specific improvements that can be made to AI systems, particularly for hardware-aware neural network design and inference optimization:
1. Hardware-Aware Neural Architecture Search (NAS)
-
Improvement: Integrate the paper's analytical CiM evaluation framework (dataflow-centric mapping, energy/latency models) directly into the NAS loop.
-
Specific Action: When searching for optimal network architectures, evaluate each candidate layer's GEMM shape (M, N, K) against the proposed CiM primitive models (Digital-6T, Analog-8T, etc.) and memory hierarchy (RF vs. SMEM). Filter out architectures that rely heavily on irregular GEMMs (e.g., M=1) or have K dimensions exceeding the CiM's reduction capability, as these show poor energy efficiency.
-
Resulting Capability: The AI system can automatically design neural networks that are not only accurate but also maximally energy-efficient and fast on a target CiM-integrated GPU, avoiding architectures that would suffer from memory bandwidth throttling or low CiM utilization.
2. Dynamic Layer-Specific Execution Scheduling
-
Improvement: Implement a runtime scheduler that uses the paper's
when to use CiM
findings to dynamically decide whether to execute a given GEMM on tensor cores or on CiM-integrated memory. -
Specific Action: For each layer during inference, compute the algorithmic reuse (operations/byte) and GEMM dimensions. If the layer has high reuse (e.g., large M, N, K) and K is within the CiM's reduction capacity, route it to CiM. If the layer is memory-bound (e.g., matrix-vector multiplication with M=1) or has a small K, route it to the standard tensor cores to avoid the CiM's higher latency and lower flexibility.
-
Resulting Capability: The AI system achieves optimal performance across heterogeneous workloads (e.g., a mix of transformer, CNN, and recommendation layers) by leveraging CiM's energy advantages where beneficial and tensor cores' flexibility elsewhere, improving overall throughput by up to 15.6× and energy efficiency by up to 3.4× compared to a static baseline.
3. Adaptive Precision and Dataflow Configuration
-
Improvement: Use the paper's dataflow mapping algorithm to dynamically configure the CiM primitives (e.g., row/column hold factors, number of active primitives) based on the current layer's GEMM shape.
-
Specific Action: For a given layer, the system calculates the optimal loop factors (e.g., maximizing weight reuse by mapping the largest possible weight matrix to CiM) and adjusts the number of active CiM primitives to achieve full utilization without over-provisioning. This prevents the performance degradation seen with skewed mappings (e.g., K:N ratio > 4) and ensures the reduction dimension K is fully processed in-situ.
-
Resulting Capability: The AI system can handle a wide range of GEMM shapes (from 16x16x16 to 8192x8192x8192) with consistent high utilization (up to 100%), avoiding the 6.6× utilization drop seen with naive heuristic mapping, and maintaining peak throughput even as input batch sizes or model dimensions change.
4. Memory Hierarchy Optimization for Inference Engines
-
Improvement: Apply the paper's
where to integrate CiM
findings to redesign the on-chip memory allocation for inference engines. -
Specific Action: Given the result that SMEM-level CiM integration (configB) outperforms RF-level integration by 10× in throughput and 0.25 TOPS/W in energy efficiency, the system should prioritize CiM integration at the largest available on-chip memory level (e.g., SMEM) over smaller levels like RF. The scheduler should also be aware that for very large workloads that don't fit in SMEM, the higher parallelism of SMEM-CiM reduces duplicate DRAM accesses.
-
Resulting Capability: The AI system can achieve higher peak performance and better energy efficiency for large-batch or large-sequence inference tasks (e.g., BERT-Large, GPT-J) by leveraging the larger capacity of SMEM for in-situ computation, reducing DRAM traffic and partial sum accesses.
5. Early-Stage Hardware-Software Co-Design Feedback
-
Improvement: Provide a quantitative feedback loop for AI model developers to estimate the hardware impact of their architectural choices (e.g., layer dimensions, precision).
-
Specific Action: The system can output a
CiM-friendliness score
for each layer, based on the paper's metrics (TOPS/W, GFLOPS, utilization). This score helps developers identify problematic layers (e.g., those with M=1 or K>1024) and refactor them (e.g., by increasing batch size, fusing layers, or adjusting embedding sizes) to improve hardware efficiency before deployment. -
Resulting Capability: AI developers can proactively design models that are inherently more efficient on CiM hardware, reducing the need for post-hoc optimization and ensuring that the final deployed system operates at near-peak energy efficiency (e.g., >1.75 TOPS/W for large regular GEMMs) rather than suffering from the 0.03 TOPS/W seen in poorly-shaped layers.
Sources
- Full Stack Optimization of Transformer Inference: a Survey
- cuDNN: Efficient Primitives for Deep Learning
- Benchmarking and modeling of analog and digital SRAM in-memory computing architectures
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- ZigZag: A Memory-Centric Rapid DNN Accelerator Design Space Exploration Framework
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Deep Residual Learning for Image Recognition
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4