Efficient transformer adaptation for analog in-memory computing via low-rank adapters
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Efficient transformer adaptation for analog in-memory computing via low-rank adapters".
Jane: The paper was written by Chen Li, Elena Ferro, Corey Lammie, Manuel Le Gallo, Irem Boybat et al. from King's College London and IBM Research Europe.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got a real mouthful of a title: “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” Jane, I’ll be honest, when I first saw that title, I had to read it twice.
Jane: You and me both, Tom. But once you unpack it, it’s actually a pretty elegant idea. So, analog in-memory computing, or AIMC, is this approach where you do math directly inside memory chips, using physical devices like phase-change memory, instead of shuttling data back and forth between memory and a processor.
Tom: Right, and that shuttling is the famous von Neumann bottleneck. It’s a huge energy and time drain, especially for big models like the transformers behind modern language AI.
Jane: Exactly. So AIMC promises to be way faster and more energy-efficient. But there’s a catch. Those analog devices are noisy. They’re not perfect digital switches. Their behavior drifts over time, and programming them is imprecise.
Tom: So you’ve got this powerful, efficient hardware, but it’s messy. And the paper’s title hints at the solution. Low-rank adapters, or LoRA. Can you break that down for our listeners?
Jane: Sure. Normally, if you want to adapt a pre-trained model to a new task, you retrain all its weights. That’s a huge undertaking. LoRA’s trick is to keep the original, big weight matrices frozen and instead train a tiny pair of new, low-rank matrices that get added to the original ones. It’s like adjusting the volume with a small knob instead of rebuilding the entire stereo system.
Tom: And the paper’s big idea is to apply that same logic to the analog hardware problem. Instead of retraining the noisy, physical weights in the memory chips, you keep them fixed and train the clean, digital LoRA adapters to compensate for the hardware’s imperfections.
Jane: That’s the core of it. The title is a bit of a mouthful, but it’s really about making these efficient analog chips practical for the flexible, adaptable AI models we actually want to use. It’s a clever marriage of old and new.
Tom: I love it. So we’ve got the title, we’ve got the gist. But I’m already wondering, does it actually work? Does a tiny digital adapter really fix a noisy analog brain? That’s what we’re going to dig into next.
Summary: Tom: So we’re back with “Efficient transformer adaptation for analog in-memory computing via low-rank adapters,” and Jane, you just hinted that the fix might be a tiny digital add-on. But the paper’s summary goes way beyond just a hunch. It’s got results.
Jane: It does, and they’re pretty compelling. They tested this idea, which they call AHWA-LoRA training, on a few different models. The first is MobileBERT, a smaller, more practical model for current hardware. And they compared their method to the old way of doing things, which was retraining the entire model to be robust to the analog noise.
Tom: And the results were on par?
Jane: Not just on par, Tom. In some cases, it was better. After simulating ten years of hardware drift, their LoRA-based method actually scored higher on the SQuAD question-answering benchmark than the full retraining approach. The F1 score was eighty-five point three six versus eighty-five point one four.
Tom: Ten years of drift, and it’s *more* accurate? That’s wild. Why would that be?
Jane: The authors have a theory. They think that by only updating the small LoRA matrices, the model stays closer to the broad, flat minima it found during pre-training. That makes it more robust to the slow, creeping changes in the analog hardware, whereas full retraining can push the model into a sharper, more fragile spot.
Tom: So it’s not just about saving effort. It’s about a fundamentally more stable way to adapt. And they didn’t stop at one model, did they?
Jane: No. They scaled it up. They showed it works on BERT-Base and BERT-Large, which are much bigger, and they found that the larger the model, the more resilient it is to the hardware noise. The performance drop over ten years was less than a single point for BERT-Large.
Tom: That’s a huge deal. It suggests that as models get bigger, this problem gets easier, not harder. But I’m a practical guy. What about the cost? This can’t be free.
Jane: That’s the beautiful part. The LoRA adapters are tiny. For MobileBERT, they’re about one point six million parameters out of a twenty-five million parameter model. And the training memory footprint is reduced by over four gigabytes compared to full retraining. It’s a massive saving.
Tom: So we’re getting equal or better accuracy, with a fraction of the trainable parameters and less memory. I’m starting to think this is too good to be true. What’s the catch? What are the actual improvements they had to make to get this to work? Let’s get into the nitty-gritty next.
Improvements: Tom: We’re back with “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” So Jane, we’ve established it works, and it’s efficient. But what are the actual improvements the paper suggests? What did they have to change to make this a reality?
Jane: The biggest improvement is in the deployment strategy. The paper proposes a hybrid architecture. The big, frozen, noisy weights stay on the analog chips. But the small, clean LoRA weights are moved to a separate, digital processor. They call these DPUs, and in their simulations, they used RISC-V based multi-core accelerators.
Tom: So you’ve got these two very different computers working together. The analog one is fast and efficient but messy, and the digital one is precise but slower. How do you make them work in harmony?
Jane: That’s the engineering challenge, and it’s all about latency balancing. You don’t want the digital processor to be a bottleneck, waiting for the analog chip to finish. So they simulated different configurations to find the sweet spot where both are busy.
Tom: And what did they find?
Jane: They found that if you process enough tokens in parallel, you can hide the latency of the LoRA computation. In their best case, the overhead of adding the LoRA adapters was only about four percent compared to a system with no adapters at all. That’s a tiny price to pay for the massive flexibility you gain.
Tom: Four percent overhead for the ability to switch tasks without reprogramming the analog chip? That sounds like a steal. And that flexibility is the other big improvement, right?
Jane: Exactly. This is the part that gets me excited. In the old way, if you wanted your analog chip to do eight different tasks, you’d need eight different models programmed onto it. That’s a huge amount of time and energy. With this method, you have one analog model, and you just swap out the tiny digital LoRA adapters for each task.
Tom: So it’s like having one universal brain, and you just change the software on the side to make it a doctor, a lawyer, or a poet.
Jane: Precisely. And they even demonstrated it on the GLUE benchmark, using one analog model to handle all eight tasks. They also showed you can adapt to new hardware conditions, like a lower-precision analog-to-digital converter, just by retraining the LoRA weights, not the whole chip.
Tom: That’s the kind of adaptability that makes this technology actually usable in the real world. But I have to ask, can this scale? We’ve talked about BERT, but what about the massive language models everyone is using now? We need to know if this holds up.
Conclusion: Tom: And we’re back for our final thoughts on “Efficient transformer adaptation for analog in-memory computing via low-rank adapters.” Jane, we’ve seen it work on smaller models, but I’m still wondering about the giants.
Jane: And that’s the most exciting part of the paper for me. They didn’t just stop at BERT. They took this idea and applied it to LLaMA three point one, an eight-billion-parameter model. That’s a model that’s hundreds of times bigger than MobileBERT.
Tom: And the LoRA adapters were still tiny?
Jane: They were about zero point five two percent of the model’s total parameters. And they used it for two different things. First, instruction tuning, where they showed it could recover a huge chunk of the performance lost when the model is deployed on noisy analog hardware. On the HellaSwag benchmark, they improved accuracy by over thirty-eight percentage points compared to the unadapted analog model.
Tom: Thirty-eight points. That’s not a small recovery. That’s bringing the model back from the dead.
Jane: And then they went even further. They used it for reinforcement learning, training the model to solve math word problems from the GSM8K dataset. They showed the analog model could go from a thirty-eight percent accuracy to over seventy percent after their training, narrowing the gap to the digital baseline significantly.
Tom: So this isn’t just a theoretical trick. It’s a practical method that works across different model sizes, different tasks, and even different training paradigms. It really feels like this could be the key to making analog hardware a mainstream reality.
Jane: I think so. The paper’s core message is that you don’t need to fight the hardware’s imperfections by retraining everything. You can work with it, using these small, adaptable digital modules to correct course. It’s a much more elegant and practical solution.
Tom: It’s a great note to end on. We’ve said goodbye to the paper, and we’re ready for the next one. Thanks for joining us, and we’ll catch you on the next episode.
Chen Li, Elena Ferro, Corey Lammie, Manuel Le Gallo, Irem Boybat, Bipin Rajendran
King's College London · IBM Research Europe
cs.AR, cs.LG
Submitted: 2026-03-21
Updated: 2026-08-17
Comments: 18 pages
Code: https://github.com/chenlicodebank/lora_on_analog_hardware
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 74/100
Key concepts
- Analog In-Memory Computing (AIMC)
- This approach performs mathematical operations directly inside memory chips using physical devices like phase-change memory. It aims to avoid the energy and time drain of moving data between separate memory and processor units.
- Low-Rank Adapters (LoRA)
- Instead of retraining all large model weights, LoRA keeps the original, large weights frozen and trains a small pair of new, low-rank matrices. These small additions compensate for hardware imperfections or adapt the model to new tasks by adjusting volume with a small knob.
- Von Neumann Bottleneck
- This is the problem where data must be shuttled back and forth between memory and a processor. This shuttling process causes significant energy consumption and time delays, which AIMC seeks to solve.
Terminology
Summary
Summary
This paper introduces Analog Hardware-Aware Low-Rank Adaptation (AHWA-LoRA) training, a novel method for efficiently adapting transformer models to Analog In-Memory Computing (AIMC) hardware. The approach addresses key limitations of conventional Analog Hardware-Aware (AHWA) training, which requires retraining the entire model and reprogramming analog devices—a process that is both time- and energy-intensive.
The proposed method consists of three main steps: (1) pre-trained transformer weights (termed meta-weights
) are directly mapped onto AIMC hardware without any training; (2) hardware constraints are simulated during forward propagation, but only lightweight LoRA (Low-Rank Adaptation) weights are updated during training; (3) the trained LoRA weights are deployed onto Digital Processing Units (DPUs), which operate in parallel with the analog computation on the fixed meta-weights. This design keeps the analog weights fixed, enabling efficient adaptation for both hardware compatibility and downstream tasks.
The paper validates AHWA-LoRA training on several benchmarks and model sizes. On the SQuAD v1.1 dataset using MobileBERT (25.3M parameters), the method achieves performance comparable to conventional AHWA training, with F1 and Exact Match (EM) scores within 1% of full AHWA training. Notably, at a 10-year conductance drift, AHWA-LoRA training outperforms conventional AHWA training (F1: 85.36 vs. 85.14; EM: 76.92 vs. 76.40). The authors attribute this improvement to the fact that LoRA updates keep the model closer to the flatter local minimum found during pretraining, improving robustness under drift.
The method significantly reduces training overhead. AHWA-LoRA training reduces trainable parameters to approximately 1.63M (a >15× reduction compared to conventional AHWA training's 24.67M) and reduces GPU memory usage by 13% (from 37.72 GB to 32.92 GB). The paper also demonstrates that hardware adaptation is inherently a low-rank problem, as adapting only 6.6% of the model parameters yields accuracy on par with full-model retraining.
For multi-task inference, the paper evaluates the method on the GLUE benchmark (8 tasks) using a single analog model mapped to one AIMC chip with 8 sets of LoRA weights (1.6M parameters each). This approach requires a total of 38.1M parameters, compared to (8 × 20.4 + 4.9)M parameters for conventional AHWA training—a more than 4-fold reduction. The method also enables on-chip task switching and adaptation to user data without reloading AIMC weights.
The paper investigates optimal LoRA allocation, finding that a rank of 8 provides the best balance between accuracy and computational overhead. Applying LoRA to all linear layers yields the highest F1 scores, while restricting LoRA to only QKV or FFN layers reduces performance. The method also supports dynamic adaptation: for example, when ADC/DAC precision is reduced from 8-bit to 6-bit, updating only the LoRA weights improves the F1 score from 60.81 to 74.23 after a 10-year drift.
Scalability studies show the method works effectively on BERT-Base (108M parameters, 1.3M LoRA weights) and BERT-Large (334M parameters, 3.5M LoRA weights). Larger models exhibit increased robustness against hardware-induced degradation: BERT-Base experiences only a 0.63-point F1 drop and BERT-Large a 0.48-point drop after 10-year drift, compared to 4 points for MobileBERT. The LoRA parameter count increases only 2× when scaling from MobileBERT to BERT-Large (12× larger), demonstrating efficient scaling.
The paper extends the method to decoder-only LLMs, specifically LLaMA 3.1 8B, for instruction tuning (using the Alpaca dataset) and reinforcement learning (using GSM8K with Group Relative Policy Optimization). For instruction tuning, the analog model before AHWA-LoRA training suffers severe degradation (e.g., HellaSwag accuracy drops from 78.91% digital baseline to 29.48%), but after AHWA-LoRA training, accuracy recovers to 67.71%, an absolute improvement of 38.23 percentage points. For reinforcement learning on GSM8K, the analog model improves from 37.98% to 70.74% accuracy after AHWA-LoRA training, reducing the analog-digital gap from 30% to 15%. The LoRA rank of 16 corresponds to only 0.52% of the model's total parameters.
Finally, the paper analyzes a hybrid hardware implementation using AIMC tiles paired with RISC-V-based Programmable Multi-Core Accelerators (PMCAs) as DPUs. By balancing AIMC tile latency with digital LoRA processing, the method achieves efficient transformer inference with only a 4% per-layer latency overhead compared to a fully AIMC implementation. The paper concludes that AHWA-LoRA training offers substantial improvements in the versatility and usability of AIMC, with minimal latency cost, and suggests that the inherent noise in AIMC could even be beneficial for reinforcement learning exploration.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and the resulting capabilities:
-
Improvement: Replace full-model retraining with LoRA-based adaptation that compensates for analog hardware noise (e.g., PCM device noise, conductance drift) while keeping the base model weights frozen.
-
Implementation: During training, inject Gaussian noise (6.7% amplitude) into the forward pass of the fixed meta-weights, but only update the LoRA parameters (rank 8). This preserves the pre-trained model's generalization while adapting to hardware imperfections.
-
Improvement: Use a single set of analog meta-weights with multiple task-specific LoRA adapters (each 1.6M parameters for MobileBERT) instead of maintaining N separate full models for N tasks.
-
Result: Achieves a >4× reduction in total parameters for 8 GLUE tasks (38.1M vs. 168.1M) while maintaining performance within 1% of full AHWA training.
-
Improvement: Enable runtime adaptation to changing hardware conditions (e.g., ADC/DAC precision reduction from 8-bit to 6-bit) by updating only the LoRA weights, not the analog array.
-
Result: Recovers F1 score from 60.81 to 74.23 after 10-year drift when hardware precision degrades.
-
Improvement: Apply AHWA-LoRA to billion-parameter LLMs (LLaMA 3.1 8B) with only 0.52% of parameters trainable, enabling training on a single 80GB GPU.
-
Result: Recovers 38.23 percentage points on HellaSwag (from 29.48% to 67.71%) and 32.76 points on GSM8K (from 37.98% to 70.74%) compared to unadapted analog models.
-
Improvement: Use GRPO with LoRA adaptation under analog noise (3.0% during RL training) to enhance reasoning capabilities.
-
Result: Achieves 70.74% on GSM8K with chain-of-thought, narrowing the analog-digital gap from 30% to 15%.
-
Low-Power Inference: Runs transformer models on analog in-memory computing hardware with only 4% per-layer latency overhead compared to fully analog implementations.
-
Long-Term Stability: Maintains accuracy within 4% of digital baseline even after 10 years of conductance drift (e.g., MobileBERT F1: 85.36 vs. 90.01 baseline).
-
Task Switching: Can switch between 8 different NLP tasks (GLUE benchmark) in real-time by swapping LoRA adapters, without reprogramming analog arrays.
-
Efficient Fine-Tuning: Adapts 8B-parameter models to new tasks with only 42M trainable parameters (0.52%), reducing GPU memory requirements by 13% compared to full AHWA training.
-
Robust Instruction Following: Maintains coherent, factual responses (e.g., correctly answering
What is a famous tall tower in Paris?
with detailed Eiffel Tower information) despite analog hardware noise. -
Reasoning Capability: Performs multi-step mathematical reasoning (GSM8K) with structured chain-of-thought outputs, achieving 70.74% accuracy under hardware constraints.
-
Lifelong Adaptation: Updates to new user data or shifting task distributions require only LoRA weight updates, not full analog array reprogramming.
-
Hardware Fault Tolerance: Automatically compensates for degradation in ADC/DAC precision, temperature-induced noise shifts, and other environmental variations through LoRA re-training.
-
Memory-Efficient Training: Trains hardware-aware models with 15× fewer trainable parameters (1.63M vs. 24.67M for MobileBERT), enabling research on larger models with limited GPU resources.
-
Transferable Methodology: The same AHWA-LoRA framework works across encoder-only (BERT) and decoder-only (LLaMA) architectures, and across supervised fine-tuning and reinforcement learning paradigms.
Abstract
Analog In-Memory Computing (AIMC) offers a promising solution to the von Neumann bottleneck. However, deploying transformer models on AIMC remains challenging due to their inherent need for flexibility and adaptability across diverse tasks. For the benefits of AIMC to be fully realized, weights of static vector-matrix multiplications must be mapped and programmed to analog devices in a weight-stationary manner. This poses two challenges for adapting a base network to hardware and downstream tasks: (i) conventional analog hardware-aware (AHWA) training requires retraining the entire model, and (ii) reprogramming analog devices is both time- and energy-intensive. To address these issues, we propose Analog Hardware-Aware Low-Rank Adaptation (AHWA-LoRA) training, a novel approach for efficiently adapting transformers to AIMC hardware. AHWA-LoRA training keeps the analog weights fixed as meta-weights and introduces lightweight external LoRA modules for both hardware and task adaptation. We validate AHWA-LoRA training on SQuAD v1.1 and the GLUE benchmark, demonstrate its scalability to larger models, and show its effectiveness in instruction tuning and reinforcement learning. We further evaluate a practical deployment scenario that balances AIMC tile latency with digital LoRA processing using optimized pipeline strategies, with RISC-V-based programmable multi-core accelerators. This hybrid architecture achieves efficient transformer inference with only a 4% per-layer overhead compared to a fully AIMC implementation.
Sources
- Carbon Emissions and Large Neural Network Training
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Analog Foundation Models
- Towards a Unified View of Parameter-Efficient Transfer Learning
- AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
- Mortal Computation: A Foundation for Biomimetic Intelligence
- The Forward-Forward Algorithm: Some Preliminary Investigations
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4