Size Matters: Foundation Model for Czech HTML documents
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Journal ref: Text, Speech, and Dialogue. TSD 2026. Lecture Notes in Computer Science(), vol 16940 141-152
DOI: 10.1007/978-3-032-37249-9_12
Project page: https://curlie.org/cs
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Creating universal, high-quality representations of web documents in high-traffic industrial environments requires models that are both performant and economic.
Terminology
Abstract
Creating universal, high-quality representations of web documents in high-traffic industrial environments requires models that are both performant and economic. Existing approaches, however, often depend on large models, overlook the structural information inherent in HTML, or are constrained by short context windows, limiting their ability to process real-world web pages. We present HTML-LM, a compact foundation model with 154 million parameters that addresses these limitations through HTML-aware training and a ModernBERT-based architecture. It was trained on 100 million web documents using multiple objectives, including masked language modeling, bag-of-words prediction, and contrastive distillation from large language models. Consequently, HTML-LM sets a new state-of-the-art for classification and regression applications in the Czech Internet domain, surpassing both larger encoders and small-sized LLMs. The model is deployed in production, processing thousands of web documents per second, and released to the community under the CC BY-NC 4.0. https://huggingface.co/Seznam/html-lm.
Sources
- DOM-LM: Learning Generalizable Representations for HTML Documents
- The Llama 3 Herd of Models
- Understanding HTML with Large Language Models
- Reasonable Effectiveness of Random Weighting: A Litmus Test for Multi-Task Learning
- Representation Learning with Contrastive Predictive Coding
- Using the Output Embedding to Improve Language Models
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- Cut Your Losses in Large-Vocabulary Language Models
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering