RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion
cs.CV, eess.IV
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Qwen3-VL Technical Report
- Emerging Properties in Self-Supervised Vision Transformers
- XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
- A Nonlocal Biharmonic Model with $\Gamma$-Convergence to Local Model and an Efficient Numerical Method
- PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
- LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
- Depth Anything 3: Recovering the Visual Space from Any Views
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
- InstantID: Zero-shot Identity-Preserving Generation in Seconds
- MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
- OmniGen: Unified Image Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models