MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum
cs.CL, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis
- Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
- Should We Still Pretrain Encoders with Masked Language Modeling?
- Distilling the Knowledge in a Neural Network
- Developing and Evaluating Tiny to Medium-Sized Turkish BERT Models
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- TurkishBERTweet: Fast and Reliable Large Language Model for Social Media Analysis
- No Language Left Behind: Scaling Human-Centered Machine Translation
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
- Representation Learning with Contrastive Predictive Coding
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering