ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
cs.CL, cs.CV
Submitted: 2026-09-08
Updated: 2026-09-08
Code: https://github.com/AshfakYeafi/Chain-DoRA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Parameter-efficient fine-tuning (PEFT) adapts large language models (LLMs) to downstream tasks while updating only a small fraction of their pretrained parameters.
Terminology
Abstract
Parameter-efficient fine-tuning (PEFT) adapts large language models (LLMs) to downstream tasks while updating only a small fraction of their pretrained parameters. Low-Rank Adaptation (LoRA) uses two trainable low-rank matrices, while Weight-Decomposed Low-Rank Adaptation (DoRA) further separates weight magnitude and direction but retains the dense LoRA-style factorization in its directional branch. We propose ChainDoRA, a weight-decomposed adaptation framework that constructs the directional low-rank factors from a connected Tensor-Train (TT) chain, where the adapter rank forms the boundary rank between input- and output-side TT contractions and an independent TT rank controls representation capacity and parameter cost. Under a controlled 15,119-example response-only adaptation setting with LLaMA-7B, ChainDoRA is evaluated against matched LoRA and DoRA baselines on seven commonsense reasoning benchmarks. ChainDoRA with TT rank 16 achieves a seven-task average accuracy of 72.30%, compared with 69.88% for LoRA and 69.39% for DoRA, while requiring only 5.35M trainable parameters versus 56.10M for LoRA and 56.98M for DoRA, corresponding to a 90.62% reduction relative to DoRA. Ablations over TT rank and adapter placement show controllable parameter-accuracy trade-offs, indicating that connected TT parameterization can substantially reduce the parameter cost of magnitude-direction adaptation while preserving, and in this setting improving, downstream reasoning performance.
Sources
- Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs
- DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
- Joint Tensor-Train Parameterization for Efficient and Expressive Low-Rank Adaptation
- Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels
- LLaMA: Open and Efficient Foundation Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Qwen2.5 Technical Report
- Mistral 7B
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- OPT: Open Pre-trained Transformer Language Models
- The Falcon Series of Open Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering