Bottom-up Modeling of Repeated Elements via Single Image Analysis-by-Synthesis
cs.CV
Submitted: 2026-09-07
Updated: 2026-09-16
Terminology
Sources
- Open-world Text-specified Object Counting
- Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames
- MONet: Unsupervised Scene Decomposition and Representation
- Structure from Duplicates: Neural Inverse Graphics from a Pile of Objects
- FastJAM: a Fast Joint Alignment Model for Images
- ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- CounTR: Transformer-based Generalised Visual Counting
- Bridging the Gap to Real-World Object-Centric Learning
- DINOv3
- Illiterate DALL-E Learns to Compose
- Unsupervised Object Discovery: A Comprehensive Survey and Unified Taxonomy
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models