WWW: What, When, Where to Compute-in-Memory
summary
The gist
This paper addresses the challenges of integrating Compute-in-Memory (CiM) paradigms into on-chip memory subsystems for efficient matrix multiplication during Machine Learning (ML) inference.
In short
The episode discusses a paper titled "WWW: What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication during Machine Learning Inference." The authors provide an analytical framework addressing what compute-in-memory primitive to use, when it benefits specific workloads like matrix multiplication shapes, and where in the memory hierarchy it should be integrated. The hosts conclude that capacity in on-chip memory is key, but compute-in-memory is not a universal solution.
Key concepts
- Compute-in-Memory (CiM)
- This refers to placing computation directly inside the memory units, such as SRAM. It aims to make AI faster and more efficient by reducing data movement between memory and processors. The paper analyzes different CiM primitives like analog versus digital designs.
- Workload Shape
- This refers to the structure of the matrix multiplication being performed, such as whether it is a large, square block or a long, skinny matrix-vector multiply. The benefits of compute-in-memory depend heavily on this shape; some shapes benefit greatly while others do not.
- Memory Hierarchy Levels
- This refers to the different storage tiers in a computer system, specifically comparing the tiny register file and the larger shared memory. The paper investigates where to place CiM units within these levels to maximize performance gains.
- Mapping Algorithm
- This is a priority-based algorithm developed by the authors that decides how to map general matrix multiplications onto CiM primitives. It prioritizes keeping weights stationary, and it was found to be faster and more effective than existing heuristic search approaches.
Terminology used across episodes
This episode discusses
- WWW: What, When, Where to Compute-in-Memory · Paper Radio
- Full Stack Optimization of Transformer Inference: a Survey
- cuDNN: Efficient Primitives for Deep Learning
- Benchmarking and modeling of analog and digital SRAM in-memory computing architectures
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- ZigZag: A Memory-Centric Rapid DNN Accelerator Design Space Exploration Framework
- LLM.int8: 8-bit Matrix Multiplication for Transformers at Scale
- Deep Residual Learning for Image Recognition
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
The paper
WWW: What, When, Where to Compute-in-Memory · Read on arXiv
Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, Kaushik Roy
Purdue University · Microsoft · Google
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "WWW: What, When, Where to Compute-in-Memory".
Jane: The paper was written by Tanvi Sharma, Mustafa Ali, Indranil Chakraborty and Kaushik Roy from Purdue University and Microsoft and Google.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone! Today we're diving into a paper that's been making the rounds on arXiv, and it's got one of those titles that just makes you stop and think: "What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication during Machine Learning Inference." Jane, I gotta say, just reading that title out loud feels like the paper is asking the universe three very specific questions.
Jane: Tom, it really does! And I love that framing, because for years we've been hearing about compute-in-memory as this magic bullet for making AI faster and more efficient. But this paper is essentially saying, "Hold on, not all compute-in-memory is created equal, and not every situation calls for it." It's like asking whether you should use a sports car or a pickup truck—it depends on what you're hauling.
Tom: Exactly! And the authors—Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, and Kaushik Roy from Purdue—they're not just theorizing. They've built an analytical framework to actually answer those three questions. The "what" is about the type of compute-in-memory primitive, the "when" is about the workload shape, and the "where" is about which level of the memory hierarchy you integrate it into.
Jane: Right, and that "where" part is so crucial. We're talking about a GPU-like architecture, with register files and shared memory. The paper is essentially saying, "If you're going to put compute inside memory, you need to know if it belongs in the tiny, super-fast register file or the bigger, slightly slower shared memory." And the answer isn't obvious until you actually do the math.
Tom: And the math is pretty compelling. They show energy efficiency improvements up to three point four times and throughput improvements up to fifteen point six times compared to a baseline tensor-core architecture. But here's the kicker—those numbers aren't universal. They depend heavily on the shape of the matrix multiplication you're doing.
Jane: That's the "when" question. Some matrix multiplications are like a big, square block—lots of reuse, lots of compute. Others are long and skinny, like a matrix-vector multiply, and those don't benefit nearly as much. In fact, for some of those skinny shapes, the compute-in-memory approach can actually be worse than just using regular cores.
Tom: So the paper isn't just cheerleading for compute-in-memory. It's giving us a roadmap for when it's actually worth the integration effort. And that's the kind of nuanced, practical insight that I think the hardware community really needs right now.
Jane: Absolutely. And it makes me wonder—what does this mean for the next generation of AI accelerators? Are we going to see chips that dynamically decide whether to use compute-in-memory or standard cores depending on the layer of the neural network?
Tom: That's a great question, and I think the paper's framework could absolutely inform that kind of adaptive design. But let's not get ahead of ourselves. We need to talk about the actual methodology they used to get these results. That's coming up next.
Summary: Tom: So Jane, we've set the stage with those three big questions—what, when, where. Now let's talk about how this paper actually goes about answering them. The core idea is that they've built a dataflow-centric way of representing compute-in-memory primitives.
Jane: Right, and I think that's a really clever move. Instead of getting bogged down in the transistor-level details of every single SRAM design out there, they abstract it into something the architecture can understand. They break a compute-in-memory primitive into smaller "CiM units," and each unit has a certain number of rows and columns it can process in parallel, and a certain number it processes sequentially.
Tom: Exactly. So you have these parameters—Rp and Cp for parallel rows and columns, and Rh and Ch for the sequential hold factors. That lets them compare very different designs, like an analog 6T SRAM primitive versus a digital 8T one, all under the same analytical umbrella.
Jane: And that's important because these primitives are wildly different. The analog ones are super energy-efficient per operation, like zero point zero nine picojoules for an eight-bit MAC, but they're slow—one hundred forty-four nanoseconds. The digital ones are faster but use more energy. Without a unified framework, you'd be comparing apples to oranges.
Tom: Right, and they also enforce iso-area constraints. So if one primitive takes up more silicon area, you get fewer of them in the same cache space. That's the fair way to compare—you're asking, "Given the same chip area, which design gives me the best performance?"
Jane: And then there's the mapping algorithm. This is where they really shine. They've written a priority-based algorithm that decides how to map a given GEMM—that's general matrix multiplication—onto the compute-in-memory primitives. The first priority is keeping the weights stationary, because that's the whole point of compute-in-memory. You want to load the weights once and then stream the inputs through them.
Tom: And they compare this to a heuristic search approach, which is what a lot of other tools use. Their algorithm is not only faster to run, but it consistently finds better mappings—up to six point six times better hardware utilization, and about one point two times better energy efficiency.
Jane: That's a big deal. It means their approach isn't just a theoretical exercise. It's a practical tool that could be used in real chip design flows to quickly evaluate whether a particular compute-in-memory design is worth pursuing.
Tom: And speaking of practical, the results they get are really interesting. They test on real workloads—ResNet50, BERT-Large, GPT-J, DLRM—and they find that the benefits vary a lot. BERT-Large, with its big, regular matrix multiplications, gets huge gains. But some of the GPT-J layers, especially during the decoding phase, have these skinny matrix-vector multiplies, and those don't benefit at all.
Jane: Right, and that brings us back to the "when" question. The paper is essentially saying, "Don't just bolt compute-in-memory onto everything. Look at your workload first." And that's such a practical, engineering-minded approach.
Tom: It really is. So we've got the framework, we've got the mapping algorithm, and we've got the workload analysis. But the real meat is in the "where" question—where in the memory hierarchy should you put this thing? That's what we're going to dig into next.
Improvements: Tom: Alright, Jane, let's get into the "where" question, because this is where the paper really earns its keep. They compared integrating compute-in-memory at two levels: the register file, which is tiny and super fast, and the shared memory, which is bigger and slower but has way more capacity.
Jane: And the results are fascinating. At the register file level, you get decent improvements—about three times better energy efficiency for BERT layers compared to a baseline tensor core. But the throughput is capped because you only have so many compute-in-memory primitives fitting in that small space.
Tom: Right, and then they look at shared memory. They actually consider two configurations. One keeps the same number of compute primitives as the register file version, just to isolate the effect of the memory level itself. The other fills the entire shared memory with compute primitives, which is a much bigger setup.
Jane: And that second configuration is where things get wild. The throughput jumps to about ten times what the register file version achieves. Because you've got so many more primitives working in parallel. But here's the twist—the energy efficiency doesn't improve as much as you'd hope.
Tom: Yeah, that surprised me too. The energy efficiency at the shared memory level is actually lower than at the register file level when you keep the same number of primitives. Because without that intermediate register file to cache data, you're going to main memory more often, and that's expensive.
Jane: But when you fill the whole shared memory with compute primitives, the energy efficiency does go up—about zero point two five TOPS/W better than the register file version. So the capacity of the memory level matters more than its position in the hierarchy.
Tom: That's a really important takeaway. It suggests that if you're designing a chip for compute-in-memory, you should focus on maximizing the amount of on-chip memory that can do computation, rather than trying to put compute in every single level.
Jane: And they also compare against the baseline tensor core architecture. The compute-in-memory at shared memory level gets up to fifteen point six times better throughput for some workloads. But for those skinny matrix-vector multiplies, it's actually worse. The baseline tensor core can be more flexible with its dataflow, while compute-in-memory is locked into a weight-stationary approach.
Tom: So the paper is really saying, "Here's the sweet spot, and here's where you should avoid." That's the kind of guidance that chip designers desperately need, because compute-in-memory is a significant investment in terms of design effort and area.
Jane: Absolutely. And it also points to future work—like, could you have a hybrid architecture that dynamically switches between compute-in-memory and standard cores depending on the layer? That would be the ultimate answer to the "when" question.
Tom: That's a fantastic vision. And it's not just academic—the paper's framework could be used to evaluate exactly that kind of hybrid design. But before we get too far into the future, let's wrap up what we've learned today.
Conclusion: Tom: Well, Jane, we've covered a lot of ground on "What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication during Machine Learning Inference." Let's pull it all together.
Jane: Absolutely. The paper gives us three clear answers. For "what," the digital 6T SRAM primitive with an adder tree design offers the best balance of throughput and energy efficiency. For "when," compute-in-memory shines with large, regular matrix multiplications—like those in BERT—but struggles with skinny matrix-vector multiplies. And for "where," the shared memory level, when fully packed with compute primitives, delivers the biggest performance gains.
Tom: And the key insight that ties it all together is that capacity matters more than hierarchy position. The more compute you can pack into on-chip memory, the better, because you're reducing those expensive trips to main memory.
Jane: Right. And they've also given us a practical mapping algorithm that's faster and more effective than existing heuristic searches. That's a tool that could be used in real chip design flows today.
Tom: The paper also highlights a sobering reality—compute-in-memory isn't a silver bullet. For memory-bound workloads with low reuse, it can actually underperform a standard tensor core. So the design guidance is nuanced, which is exactly what engineers need.
Jane: And the implications are huge. As AI models get bigger and more complex, the energy cost of moving data around becomes the dominant bottleneck. This paper gives us a roadmap for mitigating that, at least for the matrix multiplication operations that form the backbone of inference.
Tom: So, to the authors—Tanvi Sharma, Mustafa Ali, Indranil Chakraborty, and Kaushik Roy—great work. You've given the hardware community a clear, actionable framework for integrating compute-in-memory where it actually helps.
Jane: And to our listeners, if you're designing AI hardware, this paper should be on your reading list. It's not just theory; it's a practical guide with real numbers and real insights.
Tom: Alright, that's a wrap on "What, When, Where to Compute-in-Memory." Next up, we've got a paper on spiking neural networks that's been getting a lot of buzz. Until then, keep computing, and maybe do it in memory—but only when it makes sense!
Jane: See you next time, everyone!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language