MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 12 pages, 6 tables. Model weights: https://huggingface.co/moganai
License: http://creativecommons.org/licenses/by/4.0/
The gist: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed).
Terminology
Abstract
We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage: MoganColBERT-TR, a multi-vector retrieval model that, instead of compressing a query or a document into a single vector, represents it at the token level through a 768->128 projection and scores it with MaxSim late interaction. The model is not trained from scratch: the embedding model's encoder is taken as the starting point and adapted to the ColBERT objective with a single-epoch distillation phase. Training data is produced from two sources - title-to-passage pairs carved out of our own pretraining corpus in the character domain and at sentence boundaries, and two Turkish question-based retrieval sets - and is distilled from the soft scores of a cross-encoder teacher (bge-reranker-v2-m3) over one positive and seven mined negatives. We show that in hard negative mining, rank-based skipping alone is insufficient and must be combined with a group mask and a cosine ceiling. Evaluation is carried out with the official pipeline of TurkColBERT, a benchmark built for Turkish late-interaction retrieval (PLAID index, exact MaxSim), on five Turkish BEIR datasets; none of them appears in our training pool, so all five results are clean zero-shot. With 148.9M parameters, MoganColBERT-TR reaches an overall score of 37.36 (35.53 nDCG@100, 31.81 nDCG@10) averaged over the five datasets and finishes second among the five models compared: it outperforms the twice-as-large ColmmBERT-base-TR on four of five datasets and by +3.05 overall, and the benchmark's largest model by +12.30. The gap to the leading model (mLateOn) is concentrated on ArguAna-TR, the dataset with by far the longest queries.
Sources
- mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
- PyLate: Flexible Training and Retrieval for Late Interaction Models
- TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
- Towards Better Monolingual Japanese Retrievers with Multi-Vector Models
- JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources
- Should We Still Pretrain Encoders with Masked Language Modeling?
- Distilling the Knowledge in a Neural Network
- Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings
- Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
- Passage Re-ranking with BERT
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
- ColBERT-XM: A Modular Multi-Vector Representation Model for Zero-Shot Multilingual Information Retrieval
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Representation Learning with Contrastive Predictive Coding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering