LoRA-generating hypernetworks for efficient on-device LLM generative personalization
cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 19 pages, 4 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-device large language models (`LLMs'), e.g.
Terminology
Abstract
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.
Sources
- LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- Language Models are Few-Shot Learners
- Text-to-LoRA: Instant Transformer Adaption
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
- Meta-learning via Language Model In-context Tuning
- Profit: Benchmarking Personalization and Robustness Trade-off in Federated Prompt Tuning
- A Survey on In-context Learning
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- Confidential Federated Computations
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Gemma: Open Models Based on Gemini Research and Technology
- Gemma 3 Technical Report
- You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
- EigenLoRAx: Recycling Adapters to Find Principal Subspaces for Resource-Efficient Adaptation and Inference
- PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches
- Adam: A Method for Stochastic Optimization
- LongLaMP: A Benchmark for Personalized Long-form Text Generation
- Long-context LLMs Struggle with Long In-context Learning
- MEND: Meta dEmonstratioN Distillation for Efficient and Effective In-Context Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks