LoRA-GA squared: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
cs.CL, cs.AI
Submitted: 2026-08-20
Updated: 2026-09-01
Terminology
Sources
- LoRA Learns Less and Forgets Less
- OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Flora: Low-Rank Adapters Are Secretly Gradient Compressors
- LoRA+: Efficient Low Rank Adaptation of Large Models
- GoRA: Gradient-driven Adaptive Low Rank Adaptation
- LoRA: Low-Rank Adaptation of Large Language Models
- MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning
- A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA
- Adam: A Method for Stochastic Optimization
- ReLoRA: High-Rank Training Through Low-Rank Updates
- DoRA: Weight-Decomposed Low-Rank Adaptation
- AdaLomo: Low-memory Optimization with Adaptive Learning Rate
- A Survey on LoRA of Large Language Models
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
- Learning Transferable Visual Models From Natural Language Supervision
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Quantum Distance Approximation for Persistence Diagrams
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering