Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training".
Jane: Large Language Models require massive amounts of training data, and because existing public datasets often contain copyrighted or proprietary content,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show, everyone! Today we're diving into something really significant in the world of large language models with a paper titled "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training." We have a fantastic team here with us to break it down.
Jane: It sounds like this paper is tackling a big issue, Tom. They're talking about the need for data that's truly open and compliant with legal rules because the current datasets are often filled with copyrighted material. This new collection is presented as a major step in providing that ethically sourced training material for AI development.
Lu: I think this paper hits on something really interesting about the scale they're achieving here, Jane. They're aggregating about two trillion tokens from sources that are either uncopyrighted or under open licenses, and it covers a huge variety of languages and domains. That diversity is what makes this collection stand out in its size range.
Meng: From an engineering standpoint, I’m curious about the practical implications of having such a massive dataset available. How does this scale affect the actual training pipeline for new models? We need to know if handling two trillion tokens introduces any unforeseen computational bottlenecks that aren't present in existing methods.
Lalam: For me, the sheer openness and diversity of Common Corpus is incredibly important because it provides a foundation for building models that can truly serve a wider range of cultures and languages ethically. This kind of inclusive data will help ensure the AI we build isn't biased toward only one dominant source or language.
Tom: That’s a great point, Lalam, about inclusivity. So, to summarize what this paper is proposing in "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training," the core thesis is that there's a critical need for open pre-training data because existing datasets often contain proprietary content that creates legal uncertainty for researchers and developers.
Paper summary: Jane: Exactly. The authors claim they’ve introduced Common Corpus, which they describe as the largest open dataset available for LLM pre-training, totaling about two trillion tokens from uncopyrighted or openly licensed sources across many languages and domains, including a large amount of code data.
Lu: And what really makes it noteworthy is that in its size range, Common Corpus is the only one with high multilingual diversity compared to other similar collections like C4C which focuses on web pages in more than fifty languages filtered by Creative Commons Licenses.
Meng: Having that level of variety across different domains and languages sounds impressive, but I wonder how the curation process managed to maintain quality while filtering out everything that wasn't truly open or usable for training.
Lalam: The paper details a rigorous curation pipeline involving several custom tools to ensure ethical compliance and high quality, which shows a lot of thought went into making this dataset responsible. These tools address issues like text segmentation, OCR error detection, PII removal using Microsoft’s Presidio tool, and even toxicity detection with Celadon.
Tom: Those specific curation methods are what really tell us about the practical reality of building such a massive resource; it's not just dumping data together. So, moving on to the conclusion of this discussion on "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training," we've seen that this collection is presented as a solution to the problem of legal uncertainty around training data.
Jane: And what are the broader implications if we consider the authors' concluding thoughts on what this dataset means for open LLM research? They suggest that with access to such diverse and ethical data, researchers can explore language models in ways that were previously restricted by licensing concerns.
Lu: I see it as unlocking a much more robust path for multilingual pretraining because of the significant token counts available in multiple languages, with at least ten billion tokens for nine different languages mentioned in the description. This opens up new avenues for exploring how AI understands complex cultural nuances across different linguistic structures.
Paper summary: Meng: If this collection proves to be suitable for multilingual pretraining, as the paper suggests, it means we can train models that actually perform well when interacting with a global user base, not just one or two dominant ones. That has direct practical implications for deployment strategies.
Lalam: I think the biggest impact here is on the culture of AI development itself; having a massive corpus built on ethically sourced material sets a precedent for how we should approach data sourcing going forward, moving away from uncertainty toward verifiable open resources. This helps build trust in the technology.
Tom: So, to wrap up our discussion on "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training," the paper essentially argues that by creating this massive, ethically curated collection of about two trillion tokens, we can provide the necessary foundation for training LLMs without running into significant copyright issues.
Jane: And what this means in simpler terms is that we have a huge library of text and code from sources that are legally clear to use for training AI, which addresses a major hurdle in scaling up model development responsibly.
Lu: The authors show the potential for models like PleIAs 350M and PleIAs 1 point 2B performing comparably to larger models on benchmarks like MultiBLiMP, indicating that this data quality translates into actual model capability across diverse linguistic tests.
Meng: That performance comparison is what I find most encouraging from an engineering viewpoint; it suggests that the diversity of Common Corpus isn't just theoretical, it actually helps in building capable AI systems.
Lalam: And from a cultural perspective, this data supports the idea that AI can be a truly global tool because it's trained on such broad human expression across many languages and domains. It gives us material to build models that are genuinely representative of our shared knowledge base.
Tom: It sounds like we have a lot of excitement about this collection, and it really does sound like the authors are pointing toward a future where ethical data availability is central to advancing large language models. We'll be right back after the break with more research on how these massive datasets impact model efficiency.
Conclusion: Tom: So, we’ve been talking about Common Corpus, and now it's time to talk about what this paper is actually calling itself: "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training."
Jane: That title really puts the focus right where it needs to be, Tom. It immediately signals that this isn't just another big dataset; it’s about solving the ethics problem in AI training data.
Lu: I think the authors are really emphasizing that scale and compliance simultaneously, which is a tough balancing act for any data project. They’re showing how you can get massive volume without sacrificing licensing integrity across so many domains.
Meng: From an engineering standpoint, that title makes me wonder what specific mechanisms they used to ensure that "ethical" aspect was consistently applied across all those trillions of tokens and six collections.
Lalam: For me, the implication of that title is huge because it suggests a path forward where developers don't have to spend years untangling copyright issues before they can even start training models. It’s about building trust into the very foundation of the AI we create.
Tom: Exactly, Lalam! It’s about moving past the uncertainty around data sourcing so researchers can focus on actually making smarter models instead of fighting legal battles over their inputs.
Jane: And when you look at the authors, it seems like they've put a lot of work into building this pipeline, suggesting a deep understanding of both language and the legal landscape affecting AI development.
Lu: The scope implied by that title suggests they’re not just looking at one area but trying to build this comprehensive resource that spans culture, government documents, and code. That’s ambitious for any single project.
Meng: I wonder how they plan to maintain the integrity of such a vast collection as you scale it up while keeping the quality control methods robust enough for that kind of scale.
Lalam: The real impact here is that we get to see what truly diverse knowledge can be leveraged, which means the AI we build will reflect a much broader and more nuanced view of human experience than models trained on narrower data sets.
Tom: Absolutely, Lalam! It opens up a whole new dimension for how we think about what makes an AI "smart" and representative of the world.
Jane: So, to wrap up this segment, "Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training," the paper argues that solving the data sourcing problem is a prerequisite for scaling responsible AI development.
Lu: And what we’re seeing is a massive effort to build a foundation that respects both open science principles and legal realities.
Meng: This work sets a high bar, showing how to handle multi-trillion token datasets with necessary quality checks integrated from the very start.
Lalam: It gives us the raw material to create AI that genuinely understands and represents the complexity of global human knowledge, which is something I find incredibly inspiring.
Tom: And that's a big picture we need to keep in mind as we look at where this kind of data availability can take us next.
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza, Mattia Nee
PleIAs
cs.CL
Submitted: 2025-06-02
Updated: 2026-10-05
Journal ref: ICLR 2026 (Oral)
Code: https://github.com/aboSamoor/pycld2
Project page: https://microsoft.github.io/presidio
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: Large Language Models require massive amounts of training data, and because existing public datasets often contain copyrighted or proprietary content, there is a critical need for truly open
Key concepts
- Common Corpus
- The largest fully open pre-training dataset for LLMs, containing approximately 2 trillion tokens from uncopyrighted or openly licensed sources across many languages and domains.
- Data Composition
- The corpus is divided into six collections: Open Government, Open Culture, Open Science, Open Web, Open Code, and Open Semantic. These collections provide a wide variety of topics and time periods for model training.
- Curation Tools
- Rigorous tools like Segmentext for text structure analysis and OCRonos for error correction ensure high data quality. They also include Microsoft's Presidio to remove personal information (PII) and Celadon to detect various types of bias and toxicity.
Terminology
Summary
Large Language Models require massive amounts of training data, and because existing public datasets often contain copyrighted or proprietary content, there is a critical need for truly open pre-training data that complies with legal regulations. This paper introduces Common Corpus, the largest open dataset for LLM pre-training, which aggregates about two trillion tokens from uncopyrighted or openly licensed sources across a wide variety of languages and domains.
The gist
Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models by being the largest fully open pre-training dataset at about 2 trillion tokens and the only one in its size range having high multilingual diversity.
Data Composition and Scope
Common Corpus is composed of six distinct collections: Open Government, Open Culture, Open Science, Open Web, Open Code, and Open Semantic. The total collection amounts to 1998647168282 tokens. The data is diverse in terms of domains and time periods. For instance:
Open Government
This collection contains more than 407B tokens and includes Finance Commons (over 23 billion tokens) covering financial documents from sources like the Securities and Exchange Commission (SEC) reports, and Legal Commons, which includes datasets like Europarl parallel corpus and Caselaw Access Project legal cases.
Open Culture
This collection aggregates cultural heritage datasets spanning over 13 languages, including French, English, German, Spanish, Portuguese, Italian, Dutch among others. It includes large portions of data from Collections As Data (CAD) initiatives like Chronicle America and Europeana.
Data Curation and Quality Control
The development of Common Corpus involved rigorous filtering and curation processes to ensure ethical compliance and high quality. The paper highlights several custom tools developed in the Bad Data Toolbox:
-
Text Segmentation: Segmentext, a specialized language model trained for structural segmentation, is used to reconstruct editorial structure from raw character sequences, supporting segmentation of text, title, table, separator, dialog, bibliography, contact information (PII), and paratext.
-
OCR Error Detection: OCRoscope uses language detection models across rolling 7-grams to provide a document-level quality score without ground truth. OCRerrcr is a DeBERTa-v3 model fine-tuned on manually labeled errors to generate fine-grained, token-level annotations of likely erroneous tokens.
-
OCR Correction: OCRonos, a generative language model based on Llama 3 8B, is used for the correction of badly digitized text, designed to be conservative and faithful to the original source material.
-
PII Removal: Microsoft’s Presidio tool is employed to identify and replace Personally Identifiable Information (PII), including phone numbers, email addresses, IBANs, IP addresses, and URLs.
-
Toxicity Detection: Celadon is a multilingual toxicity classifier fine-tuned on DeBERTa-v3-small with five independent classification heads corresponding to dimensions like race/origin bias, gender/sexuality bias, religious bias, ability bias (physical/mental disability), and violence/abuse.
Multilingual and Code Coverage
The dataset exhibits significant multilingual diversity, with at least 10B tokens for nine languages. The token counts show a wide distribution across languages:
Top Languages:
English accounts for the largest share of tokens (968,757,721,747). Other major languages include French (275,358,437,630), German (112,127,458,251), Spanish (46.5 billion tokens), and Russian (9.4 billion tokens).
Open Code:
The Open Code collection comprises 283,227,402,898 tokens from Stack v1 and v2 data. The curation pipeline involved removing files with non-informative content (like CSV or JSON5), filtering for permissive licenses, and manual filters to discard low-quality data such as files consisting of 75% or more digits.
Model Performance Evaluation
Two small language models, PleIAs 350M and PleIAs 1.2B, were trained on filtered subsets of Common Corpus. Evaluation was conducted using benchmarks like MultiBLiMP, XStoryCloze, and XCOPA. The results show that the models perform comparably to other models of their size and show outstanding performance on MultiBLIMP,
which has more languages compared to other benchmarks. For example, the 350M model achieved a MultiBLiMP score of 0.774, comparing favorably against larger models like Gemma 3 XGLM and OLMo. The evaluation confirms that Common Corpus is "suitable for multilingual pretraining.
Improvements for AI systems
As a fastidious researcher, I have meticulously reviewed the COMMON CORPUS: THE LARGEST COLLECTION OF ETHICAL DATA FOR LLM PRE-TRAINING
paper. The core contribution is a massive, legally compliant, and highly diverse open dataset suitable for training multilingual Large Language Models (LLMs).
Here are the specific improvements and capabilities that can be achieved by integrating the Common Corpus into AI systems:
)
-
Improve LLM Pre-training Efficiency and Performance via Multilingual Diversity:
-
Enhance Model Robustness Against Data Quality Issues (OCR Noise & Bias):
-
Enable Legal/Ethical Compliance in Model Training Pipelines:
-
Develop Specialized Models for Domain-Specific Reasoning (Finance, Law, Science):
)
-
Improve LLM Pre-training Efficiency and Performance via Multilingual Diversity:
-
Enhance Model Robustness Against Data Quality Issues (OCR Noise & Bias):
-
Enable Legal/Ethical Compliance in Model Training Pipelines:
-
Develop Specialized Models for Domain-Specific Reasoning (Finance, Law, Science):
Abstract
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in pre-training datasets. In addition, it includes a large amount of code data. The diversity of data sources in terms of covered domains and time periods opens up the paths for both research and entrepreneurial needs across diverse areas of knowledge. In this paper, we present the detailed provenance of data assembling and the details of dataset filtering and curation. We train two small language models on Common Corpus and find that they perform comparably to other models of their size, indicating that our dataset is suitable for multilingual pretraining. Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models.
Sources
- Toxicity of the Commons: Curating Open-Source Pre-Training Data
- Towards Best Practices for Open Datasets for LLM Training
- Datasheet for the Pile
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
- No Language Left Behind: Scaling Human-Centered Machine Translation
- The Llama 3 Herd of Models
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Training Compute-Optimal Large Language Models
- Gemma 3 Technical Report
- DeepSeek-V3 Technical Report
- StarCoder 2 and The Stack v2: The Next Generation
- OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- A Short Survey on Small Reasoning Models: Training, Inference, Applications and Research Directions
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering