PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
cs.CV, cs.AI, cs.IR
Submitted: 2026-08-27
Updated: 2026-08-27
Terminology
Sources
- OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce Search
- OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework
- Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
- RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce
- OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment
- TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
- Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
- Generative Retrieval with Preference Optimization for E-commerce Search
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- DINOv3
- MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
- O1 Embedder: Let Retrievers Think Before Action
- Qwen3 Technical Report
- TSGR: Taobao Search Generative Retrieval
- Beyond Matching: Category-Guided Latent Intent Reasoning for Generative Retrieval in E-Commerce
- GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
- OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
- Efficient Generative Retrieval for E-commerce Search with Semantic Cluster IDs and Expert-Guided RL
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models