Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
cs.LG
Submitted: 2026-09-14
Updated: 2026-09-21
Code: https://github.com/chaofengc/IQA-PyTorch
License: http://creativecommons.org/licenses/by/4.0/
The gist: Flow-matching diffusion models have recently emerged as a strong paradigm for high-fidelity visual generation.
Terminology
Abstract
Flow-matching diffusion models have recently emerged as a strong paradigm for high-fidelity visual generation. However, their prohibitively high fine-tuning cost limits scalability to downstream tasks. While Low-Rank Adaptation (LoRA) combined with spectral initialization has demonstrated accelerated convergence and improved performance in autoregressive language models by better aligning gradient directions, we find that it fails to deliver similar gains in diffusion fine-tuning, often yielding marginal or even negative improvements over vanilla LoRA.We attribute this discrepancy to a fundamental mismatch between LoRA's low-rank parameterization and the intrinsically high-rank gradients induced by the flow-matching objective. In particular, stochastic timestep sampling introduces directionally heterogeneous gradient signals across training steps, leading to misaligned updates under low-rank constraints.To address this issue, we propose Prism-LoRA,a Principal-timestep Restricted Init via Sparse Matrix-decomposition framework that improves gradient alignment during fine-tuning. Our method consists of two key components: (i) principal timestep selection, which restricts initialization gradients to a subset of dominant timesteps to suppress effective gradient rank, and (ii) principal channel filtering, which removes task-irrelevant channels, enabling the one-step spectral initialization gradient to better align with the long-horizon optimization trajectory. Extensive experiments demonstrate that our method consistently improves both convergence speed and final performance across multiple diffusion fine-tuning benchmarks, including subject-driven generation, controllable generation, and deblurring, achieving not only performance improvement but also earlier stages of convergence over baseline LoRA and other spectral-init methods.
Sources
- Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- LoRA+: Efficient Low Rank Adaptation of Large Models
- DreamTuner: Single Image is Enough for Subject-Driven Generation
- ALLoRA: Adaptive Learning Rate Mitigates LoRA Fatal Flaws
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- LCM-LoRA: A Universal Stable-Diffusion Acceleration Module
- ControlNeXt: Powerful and Efficient Control for Image and Video Generation
- Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
- OminiControl: Minimal and Universal Control for Diffusion Transformer
- OminiControl2: Efficient Conditioning for Diffusion Transformers
- UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
- Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection
- LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning
- LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
- DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
- Delta-LoRA: Fine-Tuning High-Rank Parameters with the Delta of Low-Rank Matrices
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks