Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Even Small Reasoners Should Quote Their Sources".
Jane: Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: The authors of "Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family" are introducing two small reasoning models, Pleias-RAG-350m and Pleias-RAG-1B <ref:2504.18225#pg0,Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model>. Their central thesis is that these models can outperform smaller models, even those under four billion parameters, on standardized RAG benchmarks like HotPotQA and 2WikiMultiHopQA <ref:2504.18225#pg0,4 billion parameters, on standardized RAG benchmarks>.
Jane: They claim these models offer native support for citation and grounding using literal quotes during the inference process, rather than relying on post-hoc citation methods that other works have focused on. This is a big difference in how they handle source integration.
Lu: What really stands out from their summary is that they've trained these mid-trained models on a large synthetic dataset emulating the retrieval of multilingual open sources from the Common Corpus, which totals about two trillion tokens.
Meng: That synthetic data generation approach sounds like it’s a smart way to tackle the data frictions we’ve seen with larger models, and I'm curious if that process helps them address those practical deployment issues on-device.
Lalam: I see the focus on systematically reference grounding across leading European languages as something that really elevates the capability of these smaller models beyond just performance metrics.
Tom: They also point out that these models are competitive with popular larger models like Qwen-two point five-7B, Llama-three point one-8B, and Gemma-three-4B while specifically maintaining consistent RAG performance across those European languages <ref:2504.18225#pg0,Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B>.
Jane: The paper highlights that because of their size and ease of deployment on constrained systems, these models are positioned as valuable for trustworthy AI applications in those specific settings.
Conclusion: Tom: Thinking about the title, "Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family," it really emphasizes that accuracy and sourcing shouldn't be reserved for massive models only, which is a key point for us to consider.
Jane: I think what this paper suggests is a design philosophy where the model’s output is inherently tied to verifiable evidence through direct quotes, making the entire system more auditable from the start.
Lu: The implication here, creatively speaking, is that we could design AI agents that are fundamentally more transparent in their reasoning because they aren't just guessing or synthesizing; they are citing what they found.
Meng: From a practical standpoint, if we can guarantee systematic reference grounding with literal quotes on these smaller models, it means we reduce the uncertainty when deploying them for tasks where reliability is paramount.
Lalam: The broader impact I see is that this approach could set a new standard for how we build AI systems intended for regulated industries because it directly addresses the verifiability gap and gives us evidence-based decision support.
Tom: So, essentially, the authors are showing that you don't need billions of parameters to achieve a solid foundation in RAG and source summarization if you prioritize this native citation structure.
Jane: Exactly, and they show that this approach works consistently across different European languages too, which opens up possibilities for more inclusive AI tools globally.
Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas Hinostroza, Matthieu Delsart, Irène Girard Othman Hicheur, Anastasia Stasenko, Ivan P. Yamshchikov
cs.CL
Submitted: 2025-04-25
Updated: 2026-10-05
Code: https://github.com/huggingface/nanotron
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes.
Key concepts
- Native Citation and Grounding
- The models generate citations directly during inference using a Wikipedia-like tag syntax rather than relying on post-hoc methods. This provides higher control over how sources are presented and integrated into the final answer, ensuring evidence is intrinsically linked to the response.
- Proto-agentic Reasoning Sequence
- The models follow an iterative reasoning process covering trivial, standard, and refusal questions. This allows them to dynamically direct their own workflow based on query analysis and source evaluation, acting like a simple agent that decides the best path forward.
- Tokenizer Recycling Method
- To manage special tokens efficiently in smaller models, Pleias uses a new tokenizer variant. This method re-trains the last 19 tokens as special tokens to help the model quickly recognize instruction structures and mitigate performance drops associated with pre-allocated tokens.
- Mid-training Data Generation
- Models are mid-trained on synthetic RAG examples from the Common Corpus using a custom pipeline. This process scales training data, mitigates legal risks by reusing synthetic output, and ensures high quality through structured reasoning traces and adversarial examples.
Terminology
Summary
Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes. These models are designed to outperform SLMs below 4 billion parameters on standardized RAG benchmarks while maintaining consistent performance across leading European languages, making them valuable for trustworthy AI applications in constrained environments.
Model Design and Training
Pleias-RAG-350m and Pleias-RAG-1B are mid-trained variants of base models released in December 2024, trained exclusively on open data compiled under the Common Corpus, which totals about two trillion tokens. A critical feature is their Enhanced multilingual support for European languages with a new dedicated tokenizer with better fertility and word fidelity than Llama in French, Italian, Spanish, German, or Polish,
alongside better familiarity with source formats like PDFs.
Grounding and Verifiability
The models incorporate generated citations directly during the inference process of LLMs,
opting against post-hoc citation methods. This approach allows for higher control of source presentation and display as well as an improved integration of citation materials in the answer.
The system uses a Wikipedia-inspired syntax of tags, which is more demanding on the model side than external calling, providing higher control
over how citations are presented.
Structured Reasoning
The models employ an iterative reasoning sequence encompassing three main situations: trivial questions, standard questions, and refusal due to lack of source backing. This makes the model proto-agentic under the definition suggested by Anthropic,
dynamically directing its own processes based on query analysis and source analysis. The workflow includes a query report
with standardized outputs like answerable, trivial, reformulated, and unclear.
Tokenization Strategy
To mitigate risks associated with pre-allocated special tokens in larger models, Pleias implements a tokenizer recycling method.
This involves designing a new tokenizer variant based on the selection of the last tokens repurposed as special tokens. As shown in Figure 5, this strategy re-trains each of the last 19 tokens as special tokens, which is probably responsible for the initial significant loss drop
and allows the model to recognize instruction structures quickly.
Mid-training Data Generation
The models were trained on a large dataset of RAG examples drawn from Common Corpus with various synthetic augmentations during a process termed mid-training.
This approach focuses on the synthetic generation of training data at scale,
mitigating legal risks by using models that allow for the reuse of synthetic output. The quality and complexity are ensured through filtering steps, structured reasoning traces, adversarial examples (for diversity and complexity), and a custom synthetic pipeline that generates reasoning traces for complex answers with pre-defined steps like query analysis, source analysis, and draft.
Evaluation Metrics
Standard benchmarks used include HotPotQA, 2WikiMultiHopQA, and MuSiQue. The models are evaluated using an LLM-as-a-match (instruct version of Gemma 3 12B) with grades: “yes”, “rather yes”, “rather no”, or “no”. A key finding is that Pleias RAG models show a negligible impact
on performance loss when translated to four main European languages, demonstrating strong multilingual support. Qualitative evaluation also tests scenarios like Cross-lingual scenarios with mismatched query-context language pairs to test language adherence.
Ethical and Deployment Considerations
The architecture addresses ethical imperatives through several design choices: native citation provides Auditability
and Evidence-based decision support,
the external memory paradigm ensures Data separation
and Controlled information access,
and reliance on the Common Corpus provides Legal clarity
by training exclusively on appropriately licensed materials. This positions the models as a blueprint for AI integration in regulated industries, addressing the challenges of the verifiability gap,
the authority problem,
and the control deficit.
Future Research Directions
The current roadmap focuses on several areas: context length extension to handle longer sources, built-in support for search including generating API calls to trusted sources, personality tuning (named Pico), and reinforcement learning focused specifically on citation accuracy through structured critique. The team is also experimenting with an iterative error-informed reasoning pipeline
where model failures are accumulated as training data.
The gist: Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes. These models are designed to outperform SLMs below 4 billion parameters on standardized RAG benchmarks while maintaining consistent performance across leading European languages, making them valuable for trustworthy AI applications in constrained environments.
How it works
- The models are mid-trained on a large synthetic dataset emulating the retrieval of a wide variety of multilingual open sources from the Common Corpus.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems using the Pleias-RAG model family, based on the provided paper:
-
Disseminate a new generation of Small Language Models (SLMs) capable of high-accuracy RAG, search, and source summarization.
-
Enable native support for citation and grounding with literal quotes integrated directly into generated answers using a Wikipedia-inspired syntax of
less
tags (e.g., X">Quote). -
Implement a dynamic, proto-agentic reasoning sequence that allows models to self-determine the course of action based on query complexity and source availability. This sequence should include steps for:
-
Query reformulation for badly phrased inputs,
-
Identification of trivial questions requiring direct answering,
-
A hierarchical source analysis (reasoning re-ranker) to identify the most relevant sources, and
-
A standardized source report (extensive, basic, incomplete, or infeasible) to guide the final answer generation path.
-
Enhance multilingual support for European languages by utilizing a dedicated tokenizer with improved fertility and word fidelity compared to existing models (like Llama in French/Italian/Spanish/German/Polish).
-
Integrate a
tokenizer recycling
method to optimize special tokens, potentially improving instruction following and efficiency by repurposing the least useful tokens from the base model's vocabulary. -
Facilitate deployment on constrained infrastructure (e.g., Raspberry Pi 4) for use cases requiring local or on-device AI assistants in regulated sectors, such as legal or field support.
-
Establish a robust ethical framework for regulated industries by ensuring built-in traceability and auditability through the native citation mechanism, allowing domain experts to verify outputs against source materials.
-
Enable controlled information access via an external memory paradigm, allowing organizations to precisely define which sources the model accesses, thereby minimizing data leakage risks and maintaining clear boundaries between proprietary information and AI processing.
-
Develop tiered source reliability indicators within the citation framework, enabling governance workflows for vetting and approving sources before they are used in RAG systems.
-
Improve performance on complex, multi-hop reasoning tasks (as evidenced by 2WikiMultiHopQA and MuSiQue) by training on a synthetic mid-training dataset that incorporates:
-
Synthetic generation of training data at scale, specifically focusing on structured reasoning traces (planning, evaluation, reflection, exploration).
-
Adversarial exercises during mid-training to increase retrieval task difficulty and model resilience across source selection, shuffling, refusal design (query refusal), and language switching scenarios.
-
Produce answers in the original query language by estimating the query's language first while performing reasoning in a standardized English core sequence.
-
Implement an iterative error-informed reasoning pipeline during inference where the generated answer is critiqued by a stronger LLM or domain-specific critic to identify and correct reasoning errors before final output is presented.
Abstract
We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthetic dataset emulating the retrieval of a wide variety of multilingual open sources from the Common Corpus. They provide native support for citation and grounding with literal quotes and reintegrate multiple features associated with RAG workflows, such as query routing, query reformulation, and source reranking. Pleias-RAG-350m and Pleias-RAG-1B outperform SLMs below 4 billion parameters on standardized RAG benchmarks (HotPotQA, 2wiki) and are competitive with popular larger models, including Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B. They are the only SLMs to date maintaining consistent RAG performance across leading European languages and ensuring systematic reference grounding for statements. Due to their size and ease of deployment on constrained infrastructure and higher factuality by design, the models unlock a range of new use cases for generative AI.
Sources
- Phi-4 Technical Report
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- Large Language Models Cannot Self-Correct Reasoning Yet
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Improving Attributed Text Generation of Large Language Models via Preference Learning
- Best Practices and Lessons Learned on Synthetic Data
- Scaling Laws for Fact Memorization of Large Language Models
- Large Language Models for Power Scheduling: A User-Centric Approach
- 2 OLMo 2 Furious
- On the Capacity of Citation Generation by Large Language Models
- Qwen2.5 Technical Report
- Citekit: A Modular Toolkit for Large Language Model Citation Generation
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- InSTA: Towards Internet-Scale Training For Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering