SALD: Self-Referenced Advantage Learning for Diffusion Models
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-07
Terminology
Sources
- PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
- Spectrally-Guided Diffusion Noise Schedules
- Aligning Distributionally Robust Optimization with Practical Deep Learning Needs
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Reinforcement Learning via Self-Distillation
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
- Adaptively Point-weighting Curriculum Learning
- Flow Matching for Generative Modeling
- Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
- UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
- Learning Transferable Visual Models From Natural Language Supervision
- Learning Deep Representations of Fine-grained Visual Descriptions
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Self-Distillation Enables Continual Learning
- Score-Based Generative Modeling through Stochastic Differential Equations
- An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models
- SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
- Temporal Difference Learning for Diffusion Models
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models