H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Xiaojun Yuan, Jiangshou Hong, Guosheng Dong, Xiangzhi Wang
Alibaba Health
cs.AI, cs.LG
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 14 pages, 4 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 63/100
The gist: The paper addresses the "Retrieval Granularity Gap" in "terminology-intensive retrieval," noting that "existing representations lie at two extremes: single-vector retrievers often over-compress local
Terminology
Summary
The paper addresses the Retrieval Granularity Gap
in terminology-intensive retrieval,
noting that "existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. To bridge this gap, the authors
introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit selection with weighted MaxSim interaction."
The architecture of H+ Embedding is built on a shared bidirectional encoder
(initialized from Qwen3-0.6B-Base) that produces global, phrase-level, and lexical representations.
The model learns context-dependent retrieval units
where a linear-chain CRF (Lafferty, McCallum, and Pereira 2001) predicts a BIO label y i in O, B, I for each contextual token state.
This process ensures that the decoded units form a partition S(x) = s 1,, s m of the token positions, so every token belongs to one retrieval unit,
while retaining O tokens as singletons ensures that uncertain boundary predictions do not discard local retrieval signals.
To manage efficiency under practical constraints, the model applies budgeted unit selection with aggregated token importance for weighted phrase-level MaxSim retrieval,
where a shared importance head assigns each token a positive score, which is summed within each unit.
The training follows a Two-Stage Training
approach:
-
Stage 1: General Embedding Adaptation: The
global branch
is trainedon weakly supervised query-document pairs
usingMatryoshka representation learning (MRL)
to makeprefix widths r usable within the same encoder.
-
Stage 2: Joint Multi-Granularity Fine-Tuning: The model
jointly optimizes three retrieval heads h in G, T, L, corresponding to global, token-interaction, and lexical retrieval.
This stagecombines hard-label contrastive learning with teacher distillation within a unified objective.
Experimental results across 16 scientific, medical, and bilingual tasks
demonstrate that its phrase retrieval branch exceeds the global retrieval branch by 6.91 macro nDCG@10.
Furthermore, the phrase branch nearly matches Token while using 13.7% fewer document vectors and outperforms content-independent grouping rules under moderate vector budgets.
The authors conclude that context-dependent phrase interaction therefore provides an intermediate quality-cost point between global compression and token-level interaction for practical retrieval systems.
Improvements for AI systems
1. Adaptive Multi-Granularity Retrieval Architecture
-
The Improvement: Replace binary
single-vector vs. token-level
retrieval architectures with a unified, variable-length phrase-partitioning engine using a shared bidirectional encoder and a linear-chain CRF. -
What the improved system can do: It can dynamically switch between broad semantic matching (global) and precise technical matching (phrase-level) within a single model. It avoids the
over-compression
error of standard embeddings while avoiding the massive computational overhead of full token-level late interaction.
2. Importance-Guided Budgeted Indexing
-
The Improvement: Integrate a learned
importance head
that assigns positive scores to tokens to drive budgeted unit selection during the indexing and scoring phases. -
What the improved system can do: It can optimize storage and retrieval speed by selectively indexing only the most semantically significant phrases. This allows the system to achieve near-token-level accuracy while using significantly fewer document vectors (reducing storage/indexing costs by approximately 13.7%).
3. Terminology-Preserving Semantic Search for Specialized Domains
-
The Improvement: Implement BIO-labeling (Begin, Inside, Outside) via CRF to group subword tokens into coherent, context-dependent
retrieval units
rather than relying on static subword tokenization. -
What the improved system can do: In highly technical fields (e.g., medicine, law, or science), the system will treat complex multi-word terms (like
myocardial infarction
) as single, meaningful retrieval units. This prevents the loss of local relevance signals that occurs when standard models break technical terms into meaningless subword fragments.
4. Unified Scalable Embedding via Matryoshka-Multi-Granularity Training
-
The Improvement: Implement a two-stage training pipeline that combines Matryoshka Representation Learning (MRL) with joint multi-granularity fine-tuning (Global, Token, and Lexical heads).
-
What the improved system can do: It can serve a single model across diverse hardware environments. It can provide low-latency, low-precision
prefix-width
embeddings for edge devices/mobile search and high-precision, multi-granularity phrase embeddings for heavy-duty server-side retrieval, all without retraining the model.
Abstract
Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. This mismatch raises a natural question: can context-dependent phrases provide a useful retrieval unit between global vectors and tokens? We introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit selection with weighted MaxSim interaction. Across 16 scientific, medical, and bilingual tasks, its phrase retrieval branch exceeds the global retrieval branch by 6.91 macro nDCG@10. It also nearly matches Token while using 13.7% fewer document vectors and outperforms content-independent grouping rules under moderate vector budgets. Context-dependent phrase interaction therefore provides an intermediate quality-cost point between global compression and token-level interaction for practical retrieval systems.
Sources
- Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
- Unsupervised Dense Information Retrieval with Contrastive Learning
- MTEB: Massive Text Embedding Benchmark
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- On the Theoretical Limitations of Embedding-Based Retrieval
- C-Pack: Packed Resources For General Chinese Embeddings
- Qwen3 Technical Report
- R2MED: A Benchmark for Reasoning-Driven Medical Retrieval
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection