InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
cs.CV, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 12 pages, 6 figures. Published in IEEE Access
Journal ref: IEEE Access, vol. 14, pp. 14709-14721, 2026
DOI: 10.1109/ACCESS.2026.3655392
License: http://creativecommons.org/licenses/by/4.0/
The gist: In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection.
Terminology
Abstract
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between adjacent super-/sub-categories. However, such fixed connectors may not optimally capture the relationships within a semantic hierarchy. To address this limitation, we propose interconnected hierarchical semantic representations (InterHier), which utilize a prepended learnable context to globally guide the interpretation of prompts containing hierarchical relationships. InterHier operates in two main stages. First, it constructs a hierarchy-aware prompt by integrating super-/sub-categories and prepending a learnable context. Second, it optimizes this learnable context to align visual region embeddings and textual embeddings. InterHier consistently improves performance over methods that rely on fixed connectors and can be seamlessly integrated into existing open-vocabulary object detection models. Experiments on open-vocabulary object detection benchmarks demonstrate that InterHier achieves competitive performance against state-of-the-art methods.
Sources
- CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
- TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
- Probabilistic two-stage detection
- Open Vocabulary Object Detection with Proposal Mining and Prediction Equalization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models