Polish ModernBERT: The Long and Short of Polish Language Understanding
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/answerdotai/modernbert
License: http://creativecommons.org/licenses/by/4.0/
The gist: Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures.
Terminology
Abstract
Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures. We introduce Polish ModernBERT, a family of four Polish encoders available at Base and Large scales, each with 512-token and 8K context variants. We adapt the ModernBERT pretraining recipe through staged selection experiments and release a long-context benchmark covering legal topic classification, ideological decision-direction prediction, factual-consistency assessment over literary plot summaries, and human-rights violation assessment. Across 30 tasks, Polish ModernBERT achieves the best overall performance among the evaluated Polish encoders, reaching 83.99 and 85.11 for the Base-8K and Large-8K models, respectively. On long-context tasks, the 8K variants improve over matched Polish RoBERTa-8K baselines from 67.47 to 77.15 and from 75.88 to 78.49 at the Base and Large scales, respectively. The Base-8K model achieves this gain with 22% fewer parameters (149M vs. 190M). Efficiency measurements in representative inference setups show lower peak memory usage and latency than matched Polish RoBERTa baselines in both 512-token and 8K settings. Polish ModernBERT-8K-Base additionally achieves the best result on a Polish retrieval benchmark among the evaluated encoders below 300M parameters.
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Longformer: The Long-Document Transformer
- EuroBERT: Scaling Multilingual Encoders for European Languages
- NeoBERT: A Next-Generation BERT
- Long-Context Encoder Models for Polish Language Understanding
- Unsupervised Cross-lingual Representation Learning at Scale
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
- FastText.zip: Compressing text classification models
- SGDR: Stochastic Gradient Descent with Warm Restarts
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- 2 OLMo 2 Furious
- NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
- TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
- Multilingual E5 Text Embeddings: A Technical Report
- Pretraining Finnish ModernBERTs
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering