TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish".
Jane: TabiBERT introduces a large-scale, monolingual Turkish encoder based on the ModernBERT architecture and establishes TabiBench as a unified benchmark to address reproducibility gaps in Turkish Natural Language Processing.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we've been hearing about this paper titled "TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish," and honestly, it sounds like they are tackling some pretty big issues in Turkish Natural Language Processing. It basically claims they've built a large-scale, monolingual encoder using the ModernBERT architecture and created this TabiBench benchmark to fix the reproducibility problems everyone has been having in the field.
Jane: That makes sense, Tom; it sounds like they are trying to provide a solid foundation for Turkish research by standardizing how we compare different models. The main idea seems to be demonstrating that modern architectural advancements can actually work well on languages with complex structures like Turkish without losing performance.
Lu: I'm really intrigued by the architecture they used, especially since ModernBERT already incorporated things like rotary positional embeddings and FlashAttention technology to handle longer contexts efficiently, which is a big deal for morphologically rich languages <ref:2512.23065#pg0>. It suggests that these modern techniques aren't just for English anymore.
Meng: From an engineering standpoint, I wonder how they managed to scale the training so effectively on such a massive corpus without things becoming unstable or incredibly slow during the actual training process <ref:2512.23065#pg1>. Practical application depends on how robust that large-scale model really is in a real-world setting.
Lalam: If this TabiBERT model proves to be effective, I see it potentially improving how we build and deploy applications tailored specifically for Turkish speakers, making the AI tools much more relevant to that community <ref:2512.23065#pg0>.
Tom: Exactly; the paper is really pushing the idea that we can transfer these advanced models successfully to languages like Turkish by using a standardized setup, which is what TabiBench is designed to enforce. It’s not just about getting a higher number on one test; it’s about creating a reliable way for researchers to compare models fairly across different architectural approaches.
Jane: That standardization aspect is crucial, Tom; without that unified benchmark, comparing the progress of different Turkish NLP models feels like comparing apples and oranges because everyone uses different data splits or evaluation methods <ref:2512.23065#pg2>. They are trying to close that gap in research reproducibility by setting clear protocols for everything.
Lu: I'm also paying attention to the specific tokenization strategy they employed; the way they designed the pre-tokenization regular expression seems intentionally crafted to allocate dedicated tokens for symbols like digits and punctuation, which is smart for handling code and mathematical content <ref:2512.23065#pg0>.
Paper summary: Meng: That kind of specialized token allocation definitely speaks to practical needs; if the tokenizer handles those structural elements well, the model should perform better when dealing with things like source code or complex equations that are common in technical texts <ref:2512.23065#pg0>.
Lalam: For me, the paper's focus on academic understanding tasks is interesting because if TabiBERT can handle those long contexts well, it could fundamentally change how we process and understand complex scholarly documents in Turkish <ref:2512.23065#pg2>.
Tom: Speaking of results, the paper highlights that TabiBERT achieves an average score of seventy-seven point five eight on TabiBench, which is an improvement of one point six two absolute points over BERTurk's score of seventy-five point nine six <ref:2512.23065#pg2>. That concrete comparison shows a measurable gain in overall performance after they did all that work with the model and the benchmark together.
Jane: That improvement, even if it’s not huge, is significant when you consider how much better we can understand Turkish text now because of this research <ref:2512.23065#pg0>. It shows that incremental improvements in architecture and training can lead to tangible gains when applied correctly to a specific language domain.
Lu: The paper also mentions the context length they managed to achieve is eight thousand one hundred ninety-two tokens, which is substantially longer than the five hundred twelve tokens found in the original BERT models <ref:2512.23065#pg0>. That extended context capability, combined with FlashAttention technology, really opens up new possibilities for processing very long documents or entire books in a single go.
Meng: Eight thousand tokens is a big jump for memory usage and processing time, so I have to ask how they managed to implement those improvements in training without making the process prohibitively expensive or slow for iterative development <ref:2512.23065#pg1>. We need to know if this speed gain translates into something usable in a production environment.
Lalam: If we can process whole books or very long reports efficiently, that could really transform how we summarize and analyze large amounts of Turkish data in various industries <ref:2512.23065#pg0>.
Tom: And the paper isn't just about one model; it presents TabiBench as a unified evaluation framework with twenty-eight datasets across eight different task categories, which is what really addresses that evaluation gap they were trying to solve <ref:2512.23065#pg2>. It provides this systematic way to judge performance across classification, retrieval, and semantic similarity tasks all at once.
Paper summary: Jane: That comprehensive approach is exactly what’s needed for serious research; having a standard set of rules and data splits ensures that when researchers compare TabiBERT to BERTurk or any other model, the comparison is fair because they are all playing by the same standardized set of conditions <ref:2512.23065#pg2>.
Lu: Thinking about the implications for cultural understanding, if this model can handle complex academic texts effectively, it could potentially unlock deeper insights into Turkish literature or historical documents that were previously too cumbersome to analyze quickly <ref:2512.23065#pg2>.
Meng: I’m thinking about the real-world impact on software development; if we can reliably retrieve specific pieces of code or information from vast Turkish codebases using this architecture, it could significantly speed up development workflows for applications built in Turkish <ref:2512.23065#pg0>.
Lalam: And from a cultural perspective, if the AI can process these complex texts with high accuracy, it means we might be able to build tools that preserve and analyze the rich tapestry of Turkish language and knowledge more effectively <ref:2512.23065#pg0>.
Tom: So, to wrap up this initial look at "TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish," we've seen they introduce a powerful model based on ModernBERT and a rigorous benchmark to measure it against existing systems. It clearly shows that focused architectural refinements can yield measurable performance improvements on Turkish NLP tasks <ref:2512.23065#pg2>.
Jane: The authors’ goal was to demonstrate this transferability of modern techniques while simultaneously establishing a clear, reproducible framework for the entire field of Turkish Natural Language Processing research <ref:2512.23065#pg0>. It really sets a new baseline for what high-quality Turkish models should be capable of achieving.
Lu: I think the future work they hint at will involve pushing these context lengths even further or exploring how to adapt this framework to other morphologically rich languages, which opens up huge possibilities for cross-lingual transfer studies <ref:2512.23065#pg1>.
Meng: For us on the ground, the implication is that we have a new reference point for what’s achievable in Turkish NLP systems, which helps guide where we should allocate our engineering resources for the next generation of tools <ref:2512.23065#pg0>.
Lalam: If this research continues to develop these models and benchmarks, it suggests a future where AI tools become incredibly sophisticated at understanding and generating nuanced Turkish content across all domains <ref:2512.23065#pg0>.
Conclusion: Tom: So we've been looking at TabiBERT today, and now we're getting to the conclusion where Tom and Jane break down exactly what this paper is all about and what it means for us in Turkish NLP.
Jane: Right, Tom? This paper introduces a big model called TabiBERT which is built on the ModernBERT architecture and also sets up this new benchmark called TabiBench to make sure everyone can compare their work fairly.
Lu: I think it’s really neat how they managed to take an existing strong model like ModernBERT and adapt it for Turkish while adding these specific tweaks like RoPE embeddings and FlashAttention, which is a clever engineering move.
Meng: From my side, I'm curious about the practical scaling; they trained this thing on a huge corpus of tokens, so I wonder if that means we can actually deploy something that handles really long texts efficiently in our applications.
Lalam: For me, the real impact is how well it understands Turkish culture and complex ideas; if we can get an AI to handle academic texts with this level of context, it could open up new ways for us to analyze and preserve our heritage.
Tom: Exactly! The authors are basically showing that modern techniques from English models can be successfully applied to languages like Turkish without losing performance, which is a really important demonstration for the whole field.
Jane: And they’ve created TabiBench to give us a standard yardstick; this means researchers across different groups can use the same rules to test their new ideas, which should help make the entire research process more transparent.
Lu: I think that standardization aspect is huge because it allows us to see exactly where the current state of Turkish NLP stands right now, which is really helpful for figuring out what challenges we need to tackle next.
Meng: It’s good to hear about the benchmark setup; having fixed splits and protocols makes testing much more reliable than just running a model on some random data set.
Lalam: If this leads to better understanding of complex Turkish literature, I think it could help create tools that are way more insightful for our language and history.
Tom: It really boils down to this: TabiBERT isn't just another model; it’s a complete package that gives us both a powerful tool and the infrastructure we need to properly research and build on top of it.
Jane: So, in simple terms, they’ve given us a state-of-the-art Turkish AI encoder and provided the official testing ground so we can all measure progress against each other fairly.
Lu: The potential for future work seems really exciting because they've opened up the door to seeing how these large context capabilities translate into understanding much longer, more complex documents in Turkish.
Meng: I’m still focused on the engineering side of that long context; we need to see if this performance translates into something that doesn't require a super massive infrastructure just to run inference.
Lalam: And from a cultural view, imagine having an AI capable of deeply analyzing the nuances in our literature, that could really enrich how we interact with our own culture through technology.
Tom: It’s clear that this work lays down a solid foundation for the next generation of Turkish NLP research and development by providing both the model and the testing structure to make it happen.
Melik¸sah Türker, Asude Ebrar Kızıloglu, Onur Güngör, Susan Üsküdarlı
Department of Computer Engineering, Bogaziçi University
cs.CL
Submitted: 2025-12-28
Updated: 2026-10-06
Comments: 40 pages, 2 figures, 16 tables
Code: https://github.com/huggingface/fineweb-2
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: TabiBERT introduces a large-scale, monolingual Turkish encoder based on the ModernBERT architecture and establishes TabiBench as a unified benchmark to address reproducibility gaps in Turkish Natural
Key concepts
- ModernBERT Architecture
- This is the underlying design blueprint for TabiBERT, a powerful encoder model. It was adapted from ModernBERT to create an efficient and effective Turkish language model. This architecture allows the model to learn complex linguistic patterns in Turkish effectively.
- Rotary Positional Embeddings (RoPE)
- RoPE is a technique used in TabiBERT to help the model understand context over long sequences. It adds positional information to tokens, allowing the model to process longer texts more accurately than previous models, which is crucial for handling extended documents.
- TabiBench
- TabiBench is a unified benchmark created specifically for Turkish NLP research. It contains 28 datasets across eight diverse tasks, ensuring that researchers can compare different models fairly and rigorously against standardized protocols.
Terminology
Summary
TabiBERT introduces a large-scale, monolingual Turkish encoder based on the ModernBERT architecture and establishes TabiBench as a unified benchmark to address reproducibility gaps in Turkish Natural Language Processing. This work is significant because it demonstrates that modern architectural advances, such as those found in ModernBERT, can be effectively transferred to morphologically rich languages like Turkish without sacrificing performance, while simultaneously providing a standardized framework for rigorous comparative research in the field.
The gist: TabiBERT achieves an average score of 77.58 on TabiBench, outperforming BERTurk by 1.62 points and establishing new state-of-the-art results on five of eight task categories, with particularly strong gains on question answering (+9.55 points), code retrieval (+2.41 points), and academic understanding (+0.66 points).
Model Architecture and Training
TabiBERT is a monolingual Turkish encoder built upon the ModernBERT architecture, which incorporates several architectural innovations designed to support longer contexts and improve efficiency. Key features of the model include:
-
Rotary Positional Embeddings (RoPE) to support longer contexts.
-
FlashAttention technology for improved computational efficiency during training and inference.
-
Unpadding mechanisms to enable faster processing of long sequences, such as the 8,192-token context length achieved in TabiBERT compared to the original BERT models' 512 tokens.
The model was pretrained on a large, carefully curated corpus consisting of 86.58B tokens sampled from an 84.88B token multi-domain corpus comprising web text (73%), scientific publications (20%), source code (6%), and mathematical content (0.3%). The pretraining objective adopted is Masked Language Modeling (MLM) with a 30 per cent token masking rate, following the official ModernBERT pretraining recipe. Training was conducted over three progressive phases, scaling the recipe by 50 per cent to reach a total of 1T tokens exposed to the model.
Tokenization and Efficiency
The work utilizes a tokenizer developed for Kumru LLM, which has a vocabulary of 50,176 tokens and employs byte-pair encoding (BPE). A critical design choice in the tokenizer is its pre-tokenization regular expression (RegEx), which splits text into contiguous character spans corresponding to natural token boundaries. This design allocates dedicated tokens to symbols such as digits, punctuations, and special characters, which helps the model better understand and process code, mathematical equations, and page structure. The tokenizer was trained on a mixture of 95 per cent Turkish and 5 per cent English data.
The tokenizer exhibits specific efficiency characteristics:
[Figure 1 indicates that TabiBERT's conscious allocation of tokens to structural elements yields slightly higher fertility than BERTurk, enabling superior handling of code, mathematical content, and document structure.]
This design choice results in TabiBERT having a fertility value comparable to BERTurk despite a larger vocabulary size. In contrast, the multilingual mmBERT tokenizer exhibits 41.0 per cent higher fertility on Turkish text compared to TabiBERT’s monolingual tokenizer, reflecting greater subword fragmentation.
TabiBench: Unified Evaluation Benchmark
To address the evaluation gap
in Turkish NLP research, TabiBench was introduced as a unified evaluation framework comprising 28 datasets across eight task categories with standardized protocols, data splits, and preprocessing pipelines. The benchmark is designed around three core principles: quality over quantity, standardization (fixed train-validation-test splits), and task diversity.
The 8 major task categories covered by TabiBench include:
-
Text classification (4 datasets).
-
Token classification (4 datasets).
-
Semantic textual similarity (2 datasets).
-
Natural language inference (2 datasets).
-
Question answering (2 datasets).
-
Information retrieval (6 datasets), including TR-MTEB, BiText, and Quora-TR.
-
Code retrieval (4 datasets), including Apps-TR and CodeSearchNet-21K-subset-TR.
-
Academic understanding tasks (4 datasets), such as ThesisAbstractClassification-11K, which assesses performance on specialized scholarly text characterized by long input contexts.
Performance Evaluation and Findings
The evaluation involved systematically comparing TabiBERT against three monolingual Turkish encoder-only models: BERTurk, YTU-Cosmos-BERT, and TurkishBERTweet. The comparison was conducted after systematic hyperparameter tuning followed by fine-tuning using TabiBench to ensure fair comparisons.
Key performance findings include:
**[TabiBERT achieves an overall average of 77.58, surpassing the previous best Turkish model (BERTurk, 75.96) by 1.62 points.
Improvements for AI systems
Based on the research presented in TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish,
here are specific, actionable improvements for AI systems:
)1. Enhanced Multilingual Language Understanding via Targeted Monolingual Specialization:
Modern BERT variants (like TabiBERT) demonstrate that specialized training on monolingual data, even with architectural enhancements like Rotary Positional Embeddings (RoPE) and Flash Attention, can yield superior performance in morphologically rich languages compared to massive multilingual models like mmBERT when specific efficiency and context handling are prioritized.
-
Specific Improvement: Implement a modular
Language-Specific Encoder
pipeline where a foundational ModernBERT architecture is fine-tuned or pre-trained specifically on high-quality, domain-specific Turkish corpora (web, scientific articles). -
AI System Capability: This system will achieve state-of-the-art performance in specialized Turkish tasks like academic understanding (e.g., Thesis Abstract Classification), code retrieval, and question answering by leveraging the model's deep structural knowledge of the language and domain vocabulary, rather than relying on generalized cross-lingual representations that suffer from token fertility issues.
)2. Robust Context Handling for Long Documents:
TabiBERT supports a context length of 8,192 tokens (16x BERT's 512), and evaluation shows this extended context directly translates to superior performance in complex reasoning tasks like Question Answering and Information Retrieval when the input exceeds standard limits.
-
Specific Improvement: Integrate long-context processing capabilities (e.g., using TabiBERT's architecture) into any Turkish NLP application that handles documents longer than typical paragraph lengths, such as legal texts or scientific papers.
-
AI System Capability: The system will be capable of accurately performing document-level Question Answering and cross-document reasoning by maintaining coherence across entire long passages without truncation, which is critical for tasks like legislative analysis or complex technical troubleshooting.
)3. Domain Expertise in Specialized Fields (Code and Academia):
The TabiBERT pretraining corpus, which includes substantial amounts of scientific publications and code, allows the model to excel in tasks requiring specialized knowledge that general-purpose models often struggle with.
-
Specific Improvement: Fine-tune or fine-tune a Turkish encoder on domain-specific datasets (e.g., medical literature for clinical notes, or large code repositories) using TabiBERT as the base.
-
AI System Capability: The system will provide highly accurate code retrieval and comprehension capabilities, specifically designed to understand and reason about programming constructs in Turkish documentation or source code, as demonstrated by its strong NDCG@10 scores on specialized retrieval benchmarks.
)4. Standardized Evaluation Infrastructure for Reproducibility:
The TabiBench framework provides a standardized suite of 28 datasets across 8 task categories with fixed splits, ensuring that performance gains are attributable to architectural improvements rather than dataset-specific optimization or inconsistent preprocessing.
-
Specific Improvement: Mandate the use of the TabiBench benchmark and its standardized split generation methodology for all future Turkish NLP model development and evaluation.
-
AI System Capability: This system will offer verifiable, reproducible performance metrics across a wide spectrum of tasks (from text classification to code retrieval) that allow researchers to reliably compare architectural advancements (e.g., comparing TabiBERT vs. BERTurk) with high confidence, moving beyond anecdotal evidence from ad hoc datasets.
)5. Efficient and Resource-Aware Deployment:
TabiBERT offers significant efficiency gains, including 2.65x faster inference and a 51% smaller memory footprint compared to the large multilingual mmBERT for Turkish-focused applications.
-
Specific Improvement: Deploy Turkish NLP models using TabiBERT as the foundational encoder for production environments where computational resources (GPU/memory) are constrained, such as edge devices or low-latency APIs.
-
AI System Capability: The system will provide high-throughput, low-latency inference for real-time Turkish text analysis (e.g., social media sentiment analysis or fast search), making advanced NLP accessible in resource-constrained environments without sacrificing the performance of a larger, but less efficient, model like mmBERT.
Abstract
The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathematical content (0.3%). The model supports a context length of 8,192 tokens, sixteen times that of existing Turkish BERT models, and inherits the ModernBERT architecture's efficiency at long context. For rigorous and reproducible evaluation we introduce TabiBench, a benchmark of 27 datasets across eight task categories with standardized splits and evaluation protocols, summarized as a GLUE-style macro-average on a 0-100 scale. TabiBERT leads the Turkish models in five of eight categories and BERTurk, the previous best, in six of eight; the gains concentrate on question answering (+9.55 F1) and code retrieval (+2.41 NDCG@10), while the four short-text categories are near saturation. Its average of 77.28 exceeds BERTurk's 75.66; the multilingual mmBERT reaches 78.98 with twice the parameters and three times the training tokens, at 41% more tokens per Turkish input. We release model weights, training configurations and evaluation code as a transparent and reproducible foundation for future Turkish encoder research.
Sources
- Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- On the Cross-lingual Transferability of Monolingual Representations
- LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- Longformer: The Long-Document Transformer
- EuroBERT: Scaling Multilingual Encoders for European Languages
- Rethinking embedding coupling in pre-trained language models
- Getting the most out of your tokenizer for pre-training and domain adaptation
- PubMed 200k RCT: a Dataset for Sequential Sentence Classification in Medical Abstracts
- New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Measuring Mathematical Problem Solving With the MATH Dataset
- ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- Developing and Evaluating Tiny to Medium-Sized Turkish BERT Models
- Clinical ModernBERT: An efficient and long context encoder for biomedical text
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- TurkishBERTweet: Fast and Reliable Large Language Model for Social Media Analysis
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering