Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
cs.CV, cs.LG, cs.RO
Submitted: 2024-11-07
Updated: 2026-09-19
Comments: ICLR 2025, 28 pages, 15 figures
Project page: https://zsh2000.github.io/diff-2-in-1.github.io
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks.
Terminology
Abstract
Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing them either solely for off-the-shelf data augmentation or as mere feature extractors. In contrast to these isolated and thus sub-optimal efforts, we introduce a unified, versatile, diffusion-based framework, Diff-2-in-1, that can simultaneously handle both multi-modal data generation and dense visual perception, through a unique exploitation of the diffusion-denoising process. Within this framework, we further enhance discriminative visual perception via multi-modal generation, by utilizing the denoising network to create multi-modal data that mirror the distribution of the original training set. Importantly, Diff-2-in-1 optimizes the utilization of the created diverse and faithful data by leveraging a novel self-improving learning mechanism. Comprehensive experimental evaluations validate the effectiveness of our framework, showcasing consistent performance improvements across various discriminative backbones and high-quality multi-modal data generation characterized by both realism and usefulness. Our project website is available at https://zsh2000.github.io/diff-2-in-1.github.io/.
Sources
- ClipCap: CLIP Prefix for Image Captioning
- Monocular Depth Estimation using Diffusion Models
- Semantic Image Synthesis via Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models