Efficient transformer adaptation for analog in-memory computing via low-rank adapters
summary
In short
The episode discusses a paper on using low-rank adapters to adapt transformer models for analog in-memory computing (AIMC). The hosts explore how these small digital adapters compensate for analog hardware noise, showing improved accuracy and efficiency across various model sizes, including large language models.
Key concepts
- Analog In-Memory Computing (AIMC)
- This approach performs mathematical operations directly inside memory chips using physical devices like phase-change memory. It aims to avoid the energy and time drain of moving data between separate memory and processor units.
- Low-Rank Adapters (LoRA)
- Instead of retraining all large model weights, LoRA keeps the original, large weights frozen and trains a small pair of new, low-rank matrices. These small additions compensate for hardware imperfections or adapt the model to new tasks by adjusting volume with a small knob.
- Von Neumann Bottleneck
- This is the problem where data must be shuttled back and forth between memory and a processor. This shuttling process causes significant energy consumption and time delays, which AIMC seeks to solve.
Terminology used across episodes
This episode discusses
- Efficient transformer adaptation for analog in-memory computing via low-rank adapters · Paper Radio
- Carbon Emissions and Large Neural Network Training
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Analog Foundation Models
- Towards a Unified View of Parameter-Efficient Transfer Learning
- AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
- Mortal Computation: A Foundation for Biomimetic Intelligence
- The Forward-Forward Algorithm: Some Preliminary Investigations
The paper
Efficient transformer adaptation for analog in-memory computing via low-rank adapters · Read on arXiv
Chen Li, Elena Ferro, Corey Lammie, Manuel Le Gallo, Irem Boybat, Bipin Rajendran
King's College London · IBM Research Europe
Analog In-Memory Computing (AIMC) offers a promising solution to the von Neumann bottleneck. However, deploying transformer models on AIMC remains challenging due to their inherent need for flexibility and adaptability across diverse tasks. For the benefits of AIMC to be fully realized, weights of static vector-matrix multiplications must be mapped and programmed to analog devices in a weight-stationary manner. This poses two challenges for adapting a base network to hardware and downstream tasks: (i) conventional analog hardware-aware (AHWA) training requires retraining the entire model, and (ii) reprogramming analog devices is both time- and energy-intensive. To address these issues, we propose Analog Hardware-Aware Low-Rank Adaptation (AHWA-LoRA) training, a novel approach for efficiently adapting transformers to AIMC hardware. AHWA-LoRA training keeps the analog weights fixed as meta-weights and introduces lightweight external LoRA modules for both hardware and task adaptation. We validate AHWA-LoRA training on SQuAD v1.1 and the GLUE benchmark, demonstrate its scalability to larger models, and show its effectiveness in instruction tuning and reinforcement learning. We further evaluate a practical deployment scenario that balances AIMC tile latency with digital LoRA processing using optimized pipeline strategies, with RISC-V-based programmable multi-core accelerators. This hybrid architecture achieves efficient transformer inference with only a 4% per-layer overhead compared to a fully AIMC implementation.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Efficient transformer adaptation for analog in-memory computing via low-rank adapters".
Jane: The paper was written by Chen Li, Elena Ferro, Corey Lammie, Manuel Le Gallo, Irem Boybat et al. from King's College London and IBM Research Europe.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got a real mouthful of a title: “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” Jane, I’ll be honest, when I first saw that title, I had to read it twice.
Jane: You and me both, Tom. But once you unpack it, it’s actually a pretty elegant idea. So, analog in-memory computing, or AIMC, is this approach where you do math directly inside memory chips, using physical devices like phase-change memory, instead of shuttling data back and forth between memory and a processor.
Tom: Right, and that shuttling is the famous von Neumann bottleneck. It’s a huge energy and time drain, especially for big models like the transformers behind modern language AI.
Jane: Exactly. So AIMC promises to be way faster and more energy-efficient. But there’s a catch. Those analog devices are noisy. They’re not perfect digital switches. Their behavior drifts over time, and programming them is imprecise.
Tom: So you’ve got this powerful, efficient hardware, but it’s messy. And the paper’s title hints at the solution. Low-rank adapters, or LoRA. Can you break that down for our listeners?
Jane: Sure. Normally, if you want to adapt a pre-trained model to a new task, you retrain all its weights. That’s a huge undertaking. LoRA’s trick is to keep the original, big weight matrices frozen and instead train a tiny pair of new, low-rank matrices that get added to the original ones. It’s like adjusting the volume with a small knob instead of rebuilding the entire stereo system.
Tom: And the paper’s big idea is to apply that same logic to the analog hardware problem. Instead of retraining the noisy, physical weights in the memory chips, you keep them fixed and train the clean, digital LoRA adapters to compensate for the hardware’s imperfections.
Jane: That’s the core of it. The title is a bit of a mouthful, but it’s really about making these efficient analog chips practical for the flexible, adaptable AI models we actually want to use. It’s a clever marriage of old and new.
Tom: I love it. So we’ve got the title, we’ve got the gist. But I’m already wondering, does it actually work? Does a tiny digital adapter really fix a noisy analog brain? That’s what we’re going to dig into next.
Summary: Tom: So we’re back with “Efficient transformer adaptation for analog in-memory computing via low-rank adapters,” and Jane, you just hinted that the fix might be a tiny digital add-on. But the paper’s summary goes way beyond just a hunch. It’s got results.
Jane: It does, and they’re pretty compelling. They tested this idea, which they call AHWA-LoRA training, on a few different models. The first is MobileBERT, a smaller, more practical model for current hardware. And they compared their method to the old way of doing things, which was retraining the entire model to be robust to the analog noise.
Tom: And the results were on par?
Jane: Not just on par, Tom. In some cases, it was better. After simulating ten years of hardware drift, their LoRA-based method actually scored higher on the SQuAD question-answering benchmark than the full retraining approach. The F1 score was eighty-five point three six versus eighty-five point one four.
Tom: Ten years of drift, and it’s *more* accurate? That’s wild. Why would that be?
Jane: The authors have a theory. They think that by only updating the small LoRA matrices, the model stays closer to the broad, flat minima it found during pre-training. That makes it more robust to the slow, creeping changes in the analog hardware, whereas full retraining can push the model into a sharper, more fragile spot.
Tom: So it’s not just about saving effort. It’s about a fundamentally more stable way to adapt. And they didn’t stop at one model, did they?
Jane: No. They scaled it up. They showed it works on BERT-Base and BERT-Large, which are much bigger, and they found that the larger the model, the more resilient it is to the hardware noise. The performance drop over ten years was less than a single point for BERT-Large.
Tom: That’s a huge deal. It suggests that as models get bigger, this problem gets easier, not harder. But I’m a practical guy. What about the cost? This can’t be free.
Jane: That’s the beautiful part. The LoRA adapters are tiny. For MobileBERT, they’re about one point six million parameters out of a twenty-five million parameter model. And the training memory footprint is reduced by over four gigabytes compared to full retraining. It’s a massive saving.
Tom: So we’re getting equal or better accuracy, with a fraction of the trainable parameters and less memory. I’m starting to think this is too good to be true. What’s the catch? What are the actual improvements they had to make to get this to work? Let’s get into the nitty-gritty next.
Improvements: Tom: We’re back with “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” So Jane, we’ve established it works, and it’s efficient. But what are the actual improvements the paper suggests? What did they have to change to make this a reality?
Jane: The biggest improvement is in the deployment strategy. The paper proposes a hybrid architecture. The big, frozen, noisy weights stay on the analog chips. But the small, clean LoRA weights are moved to a separate, digital processor. They call these DPUs, and in their simulations, they used RISC-V based multi-core accelerators.
Tom: So you’ve got these two very different computers working together. The analog one is fast and efficient but messy, and the digital one is precise but slower. How do you make them work in harmony?
Jane: That’s the engineering challenge, and it’s all about latency balancing. You don’t want the digital processor to be a bottleneck, waiting for the analog chip to finish. So they simulated different configurations to find the sweet spot where both are busy.
Tom: And what did they find?
Jane: They found that if you process enough tokens in parallel, you can hide the latency of the LoRA computation. In their best case, the overhead of adding the LoRA adapters was only about four percent compared to a system with no adapters at all. That’s a tiny price to pay for the massive flexibility you gain.
Tom: Four percent overhead for the ability to switch tasks without reprogramming the analog chip? That sounds like a steal. And that flexibility is the other big improvement, right?
Jane: Exactly. This is the part that gets me excited. In the old way, if you wanted your analog chip to do eight different tasks, you’d need eight different models programmed onto it. That’s a huge amount of time and energy. With this method, you have one analog model, and you just swap out the tiny digital LoRA adapters for each task.
Tom: So it’s like having one universal brain, and you just change the software on the side to make it a doctor, a lawyer, or a poet.
Jane: Precisely. And they even demonstrated it on the GLUE benchmark, using one analog model to handle all eight tasks. They also showed you can adapt to new hardware conditions, like a lower-precision analog-to-digital converter, just by retraining the LoRA weights, not the whole chip.
Tom: That’s the kind of adaptability that makes this technology actually usable in the real world. But I have to ask, can this scale? We’ve talked about BERT, but what about the massive language models everyone is using now? We need to know if this holds up.
Conclusion: Tom: And we’re back for our final thoughts on “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” Jane, we’ve seen it work on smaller models, but I’m still wondering about the giants.
Jane: And that’s the most exciting part of the paper for me. They didn’t just stop at BERT. They took this idea and applied it to LLaMA three point one, an eight-billion-parameter model. That’s a model that’s hundreds of times bigger than MobileBERT.
Tom: And the LoRA adapters were still tiny?
Jane: They were about zero point five two percent of the model’s total parameters. And they used it for two different things. First, instruction tuning, where they showed it could recover a huge chunk of the performance lost when the model is deployed on noisy analog hardware. On the HellaSwag benchmark, they improved accuracy by over thirty-eight percentage points compared to the unadapted analog model.
Tom: Thirty-eight points. That’s not a small recovery. That’s bringing the model back from the dead.
Jane: And then they went even further. They used it for reinforcement learning, training the model to solve math word problems from the GSM8K dataset. They showed the analog model could go from a thirty-eight percent accuracy to over seventy percent after their training, narrowing the gap to the digital baseline significantly.
Tom: So this isn’t just a theoretical trick. It’s a practical method that works across different model sizes, different tasks, and even different training paradigms. It really feels like this could be the key to making analog hardware a mainstream reality.
Jane: I think so. The paper’s core message is that you don’t need to fight the hardware’s imperfections by retraining everything. You can work with it, using these small, adaptable digital modules to correct course. It’s a much more elegant and practical solution.
Tom: It’s a great note to end on. We’ve said goodbye to the paper, and we’re ready for the next one. Thanks for joining us, and we’ll catch you on the next episode.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language