Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection
cs.CV, cs.AI
Submitted: 2025-09-29
Updated: 2026-08-31
Terminology
Sources
- The Llama 3 Herd of Models
- Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
- LoRA: Low-Rank Adaptation of Large Language Models
- GPT-4o System Card
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Decoupled Weight Decay Regularization
- DRPT: Disentangled and Recurrent Prompt Tuning for Compositional Zero-Shot Learning
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- Representation Learning with Contrastive Predictive Coding
- Compositional Entailment Learning for Hyperbolic Vision-Language Models
- DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
- When and why vision-language models behave like bags-of-words, and what to do about it?
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models