LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LISA: Likelihood Score Alignment for Visual-condition Controllable Generation".
Jane: LISA proposes an effective regularization method, Likelihood Score Alignment, that explicitly aligns the intermediate features of a side network with an approximated likelihood score to enhance visual-condition controllable generation.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So folks, we're diving into a really interesting piece today called "LISA: Likelihood Score Alignment for Visual-condition Controllable Generation," and I'm honestly stoked about what this paper is showing us. Basically, they tackle this common setup where you have a main network providing an unconditional score and a side network handling the specific visual conditions, but they point out something missing there.
Jane: That’s right, Tom. The central idea of LISA seems to be revisiting that dual-branch approach through the lens of score-based generative modeling, suggesting that we should think about how those components actually contribute to the overall generation process. What’s the main claim they are making here about this setup?
Lu: They argue that because we view these models through a score perspective, the side network is implicitly supposed to provide a likelihood score, and LISA proposes an explicit regularization method to make it learn that role more directly. This addresses what the authors see as an underexplored aspect of the side branch's function in these dual-branch paradigms.
Meng: From my side of things, if you're just focusing on the efficiency of training, what exactly is this explicit alignment doing for us practically? Are we talking about a big time savings or something more subtle?
Lalam: I think this work is significant because it tackles how we structure information flow in generative models. If the side network learns to align its features with an approximated likelihood score, it suggests a much cleaner way for the model to handle specific visual conditions without needing extra external encoders.
Tom: Exactly, Lalam. And according to the summary, they propose this alignment by decomposing conditional scores into an unconditional part and a likelihood score part. It’s like giving the side network a clearer target instead of just letting it learn implicitly through the standard diffusion loss.
Jane: That decomposition, breaking down the conditional score into those two components, seems key to their entire proposal because it gives us a specific mathematical target for the side network to aim for. So, what is the core mechanism of LISA that they actually implement?
Paper summary: Lu: The LISA mechanism involves hooking features from a layer in the side network and projecting them into the score latent space using a lightweight decoder. They then construct an approximated likelihood score target and add a distance calculation between that output and the target as an extra regularization loss to the standard diffusion loss.
Meng: So, it’s adding this extra constraint during training, which is interesting because we usually try to keep training costs down. How much of a cost increase are they predicting with this approach?
Lalam: The paper suggests that this regularization can actually lead to better disentanglement of the side network's features for conditional modeling, and they claim this happens with negligible additional training cost and zero extra inference cost.
Tom: That’s a big deal if it holds true. They also reported significant gains in training convergence, stating that LISA can achieve over two point seven eight times faster convergence, similar to what was seen in ControlNet experiments.
Jane: Faster convergence is always a big win for development cycles, Tom. Beyond the speed of training, what about the quality of the final outputs when we actually generate images or videos using this LISA method?
Lu: The experiments show that LISA consistently improves the final synthetic results across various image and video tasks. They specifically highlight that it can improve metrics like FVD for video generation and PCK for pose-conditioned generation.
Meng: So, if I were looking at a practical application, say generating complex segmentation maps, would this method hold up as well as they claim across different model architectures?
Lalam: The generalization seems quite robust; the effectiveness was verified across a wide variety of models and tasks including Diffusion/Flow models like U-Net and DiT. It even shows good generalization to diffusion transformers trained with flow matching objectives.
Tom: That's impressive breadth for a regularization technique; it’s not just stuck on one specific architecture or task. I wonder what the authors think about the balance between different losses they are optimizing simultaneously?
Jane: The objective function they use is a joint minimization of both the standard diffusion loss and this new LISA regularization loss, represented as "min (Lmain + λLLISA)". That weighting parameter lambda seems crucial to tuning the balance between the two goals.
Paper summary: Lu: The ablation studies confirmed that an alignment depth of five and a loss weight of lambda = zero point two provided the best balance, leading to the highest PCK in pose-conditioned generation. That specific tuning really shows how sensitive the process is to those hyperparameters.
Meng: From an engineering standpoint, having a clear guideline on these optimal settings helps tremendously when we're trying to deploy something that needs to be efficient and reliable, which is what they emphasize in their conclusion.
Lalam: I think the implications for the wider AI culture are that we can start thinking about how we explicitly define the roles of different sub-networks within a larger generative system, rather than relying purely on implicit learning through massive data exposure.
Tom: It really shifts the way we view these dual-branch setups; it moves them from a black box configuration to something where we have defined roles that can be explicitly regularized for better control. So, looking ahead, what’s the big picture impact of LISA on future visual generation?
Jane: It suggests a path toward more controllable and structurally precise generation because the side network is now being steered toward an explicit likelihood score alignment rather than just learning whatever features it needs to get the basic conditional output.
Lu: I see this method potentially opening doors for even more complex, multi-condition generation where we need fine-grained spatial control, like generating precise depth maps or segmentation masks with high fidelity.
Meng: If it truly handles multiple visual conditions well, that means we could deploy these models in environments where precise scene reconstruction is necessary for applications beyond just pretty pictures. It’s about reliability in complex scenes.
Lalam: For me, the most important implication is how this technique can help build more nuanced AI systems where different parts of the architecture have clearly defined responsibilities, which will make the overall system much more understandable and trustworthy.
Tom: It sounds like LISA is making these models not just better at generating things, but better at controlling them precisely by clarifying what each component is supposed to be doing during the training process. We’ll keep watching this space for how these ideas translate into real-world deployment.
Conclusion: Tom: So, to wrap up this discussion on LISA, we're looking at how this new method aims to make visual generation more controllable by explicitly aligning side network features with a likelihood score.
Jane: That’s right, Tom; the paper focuses on a regularization technique that gives the model a clearer path for handling specific visual conditions during training.
Lu: The authors essentially tackle the dual-branch paradigm by showing how to decompose the conditional information into parts that can be regularized separately, which is really creative thinking.
Meng: From an engineering standpoint, this means we’re adding a targeted loss function to guide the side network's internal workings instead of just letting it learn blindly.
Lalam: I think what this paper really means is that we are building generative systems where the components have clearly defined responsibilities during training, which can make future AI development much more structured and reliable for us.
Tom: Exactly, Lalam; having those defined roles helps us understand exactly how the system is working under the hood.
Jane: And this approach doesn't just focus on getting pretty pictures; it’s fundamentally about improving the control we have over what the AI generates based on specific inputs.
Lu: I think this has huge implications for how we can use these models in real-world scenarios where precise spatial control, like generating accurate depth maps or segmentation masks, is absolutely critical.
Meng: If it can handle multiple visual conditions well, that opens up new avenues for deploying these models in environments that need high fidelity scene reconstruction.
Lalam: The biggest cultural impact I see is how this pushes us toward designing AI architectures where the roles of different parts are explicit, which will make the entire ecosystem more transparent.
Tom: So, we've seen how LISA works and what its immediate wins are for training speed and quality; now we need to look at what this means for the future direction of visual AI research.
Yanghao Wang, Hongxu Chen, Jiazhen Liu, Zhenqi He, Rui Liu, Zhen Wang
The Hong Kong University of Science and Technology 2Huawei Research
cs.CV
Submitted: 2026-06-25
Updated: 2026-09-28
Importance score: 83/100
The gist: LISA proposes an effective regularization method, Likelihood Score Alignment, that explicitly aligns the intermediate features of a side network with an approximated likelihood score to enhance
Key concepts
- Dual-Branch Paradigm
- This is a common setup where two networks work together: one main network provides an unconditional score, and a side network implicitly contributes a likelihood score. LISA analyzes this structure to show the side network lacks explicit regularization for its intended likelihood role.
- Score-Based Generative Modeling
- This framework views generation through the lens of scores—mathematical gradients that indicate the direction to improve an image or video. The main network provides one type of score, while LISA focuses on how the side network learns a second, related score component.
- Likelihood Score Decomposition
- Bayes' rule allows conditional scores to be split into two parts: an unconditional score and a likelihood score. LISA leverages this decomposition by training the side network to learn the residual gap, which corresponds directly to this crucial likelihood score component.
- LISA Mechanism
- LISA enforces alignment by projecting intermediate features from the side network into a latent space using a lightweight decoder. This output is then compared against an approximated target likelihood score, and the difference acts as an additional regularization loss.
Terminology
Summary
LISA proposes an effective regularization method, Likelihood Score Alignment, that explicitly aligns the intermediate features of a side network with an approximated likelihood score to enhance visual-condition controllable generation. This technique addresses the underexplored role of side branches in dual-branch paradigms by decomposing conditional scores into unconditional and likelihood scores.
The gist
LISA is an effective regularization method that explicitly aligns the intermediate feature of the side network with an approximated likelihood score.
The core concept
The paper revisits the dual-branch paradigm by viewing it through score-based generative modeling, where the main network provides an unconditional score and the side network implicitly contributes a likelihood score. The proposal is motivated by Bayes’ rule, which decomposes the conditional score into two components: the unconditional score ∇xt log pt(xt) and the likelihood score ∇xt log pt(cxt).
The side network's role is to learn this residual gap, which corresponds to the likelihood score.
The LISA mechanism
LISA introduces an explicit constraint by aligning the side network's intermediate features with an approximated likelihood score. This process involves several steps:
-
Hooking features from a designated layer of the side network and projecting them into the score latent space using a
lightweight decoder.
-
Constructing an
approximated likelihood score target.
-
Calculating the distance between the decoder’s output and this target as an additional regularization loss, which is added to the standard diffusion loss.
Optimization objective
The final optimization objective jointly minimizes both the standard diffusion loss and LISA's regularization loss: Overall optimization objective: min (Lmain + λLLISA).
This setup allows for the joint optimization of the side network and decoder with both losses. During inference, the auxiliary decoder is discarded, using only the trained side network for generation.
Key findings and benefits
Experiments across various image/video tasks and architectures demonstrated that LISA consistently accelerates training convergence
and improves final synthetic results.
Specifically, LISA can achieve "> 2.78× faster convergence (e.g., as in ControlNet). Furthermore, the regularization encourages the side network’s features to be
more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost." Ablation studies confirm that an alignment depth of 5 and a loss weight λ = 0.2 achieve the best balance, yielding the highest PCK (Percentage of Correct Keypoints) in pose-conditioned generation. LISA also demonstrates stronger compositional generation ability
when handling multiple visual conditions.
Generalization
The effectiveness of LISA was verified across various models and tasks, including Diffusion/Flow models (U-Net, DiT, OT-FM) and image/video tasks such as Pose maps, Segmentation maps, Low-resolution images, and Depth maps. The results show that LISA generalize well to diffusion transformers trained with flow matching objectives
and enhances performance in controllable video generation by significantly improving metrics like FVD (Frechet Video Distance) and PCK. This indicates that LISA is an effective and extensible solution for visual-condition controllable generation.
Contributions
The main contributions include:
-
Analyzing the mainstream dual-branch paradigm from a novel perspective, revealing that the side network
lacks explicit regularization for its intended role.
-
Proposing an effective likelihood alignment method, LISA, which regularizes the side network’s intermediate output with an approximated likelihood score.
-
Demonstrating significant and consistent gains on both training convergence and synthesis quality across various architectures and tasks without requiring external semantic encoders.
Conclusion
LISA significantly accelerates training and bootstraps better synthetic results by explicitly decomposing roles within the dual-branch paradigm, making the side network learn its intended likelihood-score role more explicitly and efficiently.
LISA is shown to be efficient and practical for deployment
as it introduces only a negligible number of additional parameters.
References
[1] Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schonlieb, and Christian Etmann. Conditional image generation with score-based diffusion models. arXiv preprint arXiv:2111.13606, 2021.
[2] Andreas Blattmann et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
[3] Zhe Cao et al. Realtime multi-person 2d pose estimation using part affinity fields. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7291–7299, 2017.
[4] Yisol Choi et al. Controllable human image generation with personalized multi-garments. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp.
Improvements for AI systems
Here are the specific improvements that can be made to existing AI generation systems based on the LISA framework, and what those improved systems can achieve:
-
Dominance in Visual-Condition Controllable Generation:
-
Accelerated Training Convergence: Achieve significantly faster training convergence (e.g., > 2.78× as seen in ControlNet) while maintaining or improving final synthetic quality (FID, CLIP scores).
-
Enhanced Condition Fidelity and Structure Following: Produce images/videos with higher fidelity to the input conditions (e.g., better pose consistency, accurate segmentation maps) because the side network is explicitly regularized to learn a likelihood score.
-
Improved Feature Disentanglement for Compositional Control: The side network's intermediate features become more disentangled, allowing for superior compositional control—meaning multiple visual conditions (e.g., pose + segmentation) can be integrated more effectively without interference between the control signals.
-
Reduced Training Overhead and Inference Cost: The method introduces only a negligible additional training cost and zero extra inference cost because the lightweight decoder is dropped during inference, allowing for deployment of high-quality controllable generation models with minimal latency impact.
-
Architecture Agnosticism: The LISA framework generalizes effectively across various architectures, including U-Net/DiT and Flow Matching models (like OT-FM), ensuring its effectiveness is not limited to a single model type.
-
Versatility Across Modalities: The system can be applied consistently across diverse image tasks (pose, segmentation, depth maps) and video generation tasks (pose videos), showing consistent gains in all relevant metrics.
This improved AI system can perform the following specific tasks:
-
Generate highly accurate images or videos conditioned on complex spatial inputs like precise 3D pose maps, depth maps, or semantic segmentation masks, achieving state-of-the-art structural alignment (e.g., PCK improvement from 19.38% to 83.02% in pose generation).
-
Enable rapid prototyping and fine-tuning of controllable generative models by drastically reducing the necessary training iterations required to achieve high perceptual quality (lower FID) compared to baseline methods, leading to faster model deployment cycles.
-
Create complex, multi-faceted scenes or videos where multiple visual constraints must coexist coherently (e.g., generating a video where the subject's pose is controlled by one map and the background segmentation is controlled by another), due to the side network's improved disentanglement capabilities.
-
Deploy robust, efficient conditional generation pipelines in real-time applications, as the inference cost remains comparable to standard diffusion models without needing a heavy auxiliary network during sampling.
Sources
- Conditional Image Generation with Score-Based Diffusion Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Classifier-Free Diffusion Guidance
- JOG3R: Towards 3D-Consistent Video Generators
- Composer: Creative and Controllable Image Synthesis with Composable Conditions
- Gotta Go Fast When Generating Data with Score-Based Models
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Decoupled Weight Decay Regularization
- DINOv2: Learning Robust Visual Features without Supervision
- ControlNeXt: Powerful and Efficient Control for Image and Video Generation
- Score-Based Generative Modeling through Stochastic Differential Equations
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- Wan: Open and Advanced Large-Scale Video Generative Models
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- DwNet: Dense warp-based network for pose-guided human video generation
- Endless World: Real-Time 3D-Aware Long Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models