ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving
cs.LG, cs.DC
Submitted: 2025-05-20
Updated: 2026-09-19
Code: https://github.com/huggingface/peft
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Multi-LoRA (Low-Rank Adaptation) serving allows many specialized LLM variants to share the same base model by attaching lightweight adapters.
Terminology
Abstract
Multi-LoRA (Low-Rank Adaptation) serving allows many specialized LLM variants to share the same base model by attaching lightweight adapters. This makes it attractive for serving large catalogs of domain-, tenant-, and task-specific models. However, existing deployments waste resources on rarely used adapters, while directly running LoRA inference on serverless platforms introduces high memory overhead, repeated cold starts, and poor scheduling decisions. This paper presents ServerlessLoRA, a serverless multi-LoRA serving system that separates serving state into shared backbone state, reusable variant warm state, and request-private state. ServerlessLoRA loads each backbone once, shares it read-only across isolated LoRA functions, batches selected backbone operations, selectively warms variant state, and routes requests based on state locality and GPU contention. Evaluated on industrial traces, ServerlessLoRA reduces time-to-first-token by up to 92.3% against serverless baselines and achieves 1.66-3.01 times and 2.20-3.26 times higher latency-cost efficiency than vLLM-LoRA and dLoRA, respectively.
Sources
- Training Verifiers to Solve Math Word Problems
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- {\lambda}Scale: Enabling Fast Scaling for Serverless Large Language Model Inference
- Multi-LoRA Composition for Image Generation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks