Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining
cs.CL, cs.LG
Submitted: 2026-06-20
Updated: 2026-09-08
Comments: Code, models, and data: https://github.com/doctolib-lab/doctobert
Code: https://github.com/doctolib-lab/doctobert
License: http://creativecommons.org/licenses/by/4.0/
The gist: Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining.
Terminology
Abstract
Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in non-English clinical settings. Whether web-scale data curation also benefits encoder Masked Language Modeling (MLM) in a dense-terminology domain remains an open question. To address this, we introduce two complementary levers. Medical-term density filtering selects documents rich in medical terms. Signal-amplifying rephrasing uses an LLM to rewrite documents into denser variants with broader entity contexts. We instantiate the recipe on French medical NLP. The medical-term density filter outperforms the widely-used educational quality filter on downstream medical tasks, and the two complement each other. Signal-amplifying rephrasing alone improves on raw web data, and mixing it with filtered web data produces the largest gain. The recipe yields FineMed, a French medical pretraining corpus, and DoctoBERT, a state-of-the-art French medical encoder family evaluated on both the public benchmark DrBenchmark and a proprietary clinical Named Entity Recognition (NER) task.
Sources
- Publicly Available Clinical BERT Embeddings
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- SciBERT: A Pretrained Language Model for Scientific Text
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- EuroBERT: Scaling Multilingual Encoders for European Languages
- NeoBERT: A Next-Generation BERT
- Scaling Laws and Interpretability of Learning from Repeated Data
- Language Models are Few-Shot Learners
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Bioformer: an efficient transformer language model for biomedical text mining
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains
- DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
- Should We Still Pretrain Encoders with Masked Language Modeling?
- Clinical ModernBERT: An efficient and long context encoder for biomedical text
- Reformulation for Pretraining Data Augmentation
- PMI-Masking: Principled masking of correlated spans
- DataComp-LM: In search of the next generation of training sets for language models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering