Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
cs.CL, cs.AI
Submitted: 2026-09-07
Updated: 2026-09-18
Comments: 16 pages, 9 figures, EMNLP 2026 Main
Code: https://github.com/AMALIA-LLM/datatrove-amalia
License: http://creativecommons.org/licenses/by/4.0/
The gist: Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale.
Terminology
Abstract
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.
Sources
- SemDeDup: Data-efficient learning at web-scale through semantic deduplication
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- \"UberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset
- The Llama 3 Herd of Models
- EuroLLM-9B: Technical Report
- 2 OLMo 2 Furious
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering