PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining
cs.CV
Submitted: 2026-09-19
Updated: 2026-09-19
Code: https://github.com/black-forest-labs/flux
Terminology
Sources
- Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
- OmniPSD: Layered PSD Generation with Diffusion Transformer
- OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
- Qwen-Image Technical Report
- Simpler Diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion
- PixelFlow: Pixel-Space Generative Models with Flow
- PixNerd: Pixel Neural Field Diffusion
- Back to Basics: Let Denoising Generative Models Denoise
- One-step Latent-free Image Generation with Pixel Mean Flows
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- Text2Layer: Layered Image Generation using Latent Diffusion Model
- PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
- COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- SAM 2: Segment Anything in Images and Videos
- AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
- Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
- Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models