H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases

arXiv:2608.00065 · cs.AI, cs.LG · Submitted 2026-08-07 · Read on arXiv

Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Xiaojun Yuan, Jiangshou Hong, Guosheng Dong, Xiangzhi Wang

Alibaba Health

cs.AI, cs.LG

Submitted: 2026-08-07

Updated: 2026-08-10

Comments: 14 pages, 4 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: The paper addresses the "Retrieval Granularity Gap" in "terminology-intensive retrieval," noting that "existing representations lie at two extremes: single-vector retrievers often over-compress local

Terminology

Summary

The paper addresses the Retrieval Granularity Gap in terminology-intensive retrieval, noting that "existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. To bridge this gap, the authors introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit selection with weighted MaxSim interaction."

The architecture of H+ Embedding is built on a shared bidirectional encoder (initialized from Qwen3-0.6B-Base) that produces global, phrase-level, and lexical representations. The model learns context-dependent retrieval units where a linear-chain CRF (Lafferty, McCallum, and Pereira 2001) predicts a BIO label y i in O, B, I for each contextual token state. This process ensures that the decoded units form a partition S(x) = s 1,, s m of the token positions, so every token belongs to one retrieval unit, while retaining O tokens as singletons ensures that uncertain boundary predictions do not discard local retrieval signals. To manage efficiency under practical constraints, the model applies budgeted unit selection with aggregated token importance for weighted phrase-level MaxSim retrieval, where a shared importance head assigns each token a positive score, which is summed within each unit.

The training follows a Two-Stage Training approach:

  • Stage 1: General Embedding Adaptation: The global branch is trained on weakly supervised query-document pairs using Matryoshka representation learning (MRL) to make prefix widths r usable within the same encoder.

  • Stage 2: Joint Multi-Granularity Fine-Tuning: The model jointly optimizes three retrieval heads h in G, T, L, corresponding to global, token-interaction, and lexical retrieval. This stage combines hard-label contrastive learning with teacher distillation within a unified objective.

Experimental results across 16 scientific, medical, and bilingual tasks demonstrate that its phrase retrieval branch exceeds the global retrieval branch by 6.91 macro nDCG@10. Furthermore, the phrase branch nearly matches Token while using 13.7% fewer document vectors and outperforms content-independent grouping rules under moderate vector budgets. The authors conclude that context-dependent phrase interaction therefore provides an intermediate quality-cost point between global compression and token-level interaction for practical retrieval systems.

Improvements for AI systems

1. Adaptive Multi-Granularity Retrieval Architecture

  • The Improvement: Replace binary single-vector vs. token-level retrieval architectures with a unified, variable-length phrase-partitioning engine using a shared bidirectional encoder and a linear-chain CRF.

  • What the improved system can do: It can dynamically switch between broad semantic matching (global) and precise technical matching (phrase-level) within a single model. It avoids the over-compression error of standard embeddings while avoiding the massive computational overhead of full token-level late interaction.

2. Importance-Guided Budgeted Indexing

  • The Improvement: Integrate a learned importance head that assigns positive scores to tokens to drive budgeted unit selection during the indexing and scoring phases.

  • What the improved system can do: It can optimize storage and retrieval speed by selectively indexing only the most semantically significant phrases. This allows the system to achieve near-token-level accuracy while using significantly fewer document vectors (reducing storage/indexing costs by approximately 13.7%).

3. Terminology-Preserving Semantic Search for Specialized Domains

  • The Improvement: Implement BIO-labeling (Begin, Inside, Outside) via CRF to group subword tokens into coherent, context-dependent retrieval units rather than relying on static subword tokenization.

  • What the improved system can do: In highly technical fields (e.g., medicine, law, or science), the system will treat complex multi-word terms (like myocardial infarction) as single, meaningful retrieval units. This prevents the loss of local relevance signals that occurs when standard models break technical terms into meaningless subword fragments.

4. Unified Scalable Embedding via Matryoshka-Multi-Granularity Training

  • The Improvement: Implement a two-stage training pipeline that combines Matryoshka Representation Learning (MRL) with joint multi-granularity fine-tuning (Global, Token, and Lexical heads).

  • What the improved system can do: It can serve a single model across diverse hardware environments. It can provide low-latency, low-precision prefix-width embeddings for edge devices/mobile search and high-precision, multi-granularity phrase embeddings for heavy-duty server-side retrieval, all without retraining the model.

Abstract

Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. This mismatch raises a natural question: can context-dependent phrases provide a useful retrieval unit between global vectors and tokens? We introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length phrase partitions, preserves uncovered tokens as singletons, and applies importance-guided unit selection with weighted MaxSim interaction. Across 16 scientific, medical, and bilingual tasks, its phrase retrieval branch exceeds the global retrieval branch by 6.91 macro nDCG@10. It also nearly matches Token while using 13.7% fewer document vectors and outperforms content-independent grouping rules under moderate vector budgets. Context-dependent phrase interaction therefore provides an intermediate quality-cost point between global compression and token-level interaction for practical retrieval systems.

Sources

Related papers