Preserve-and-Compose Training for Composed Image Retrieval
cs.CV
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/miccunifi/SEARLE
Terminology
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
- MovieNet: A Holistic Dataset for Movie Understanding
- Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
- Vision-by-Language for Training-Free Compositional Image Retrieval
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
- Qwen3 Technical Report
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval
- Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval
- GeneCIS: A Benchmark for General Conditional Image Similarity
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
- Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models