Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

summary

Video file (mp4)

The gist

Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes.

In short

Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and summarization that natively support citations with literal quotes. These models are trained on a massive open dataset, focusing on European languages, and use structured reasoning to ensure verifiable answers. They aim to outperform smaller models in RAG benchmarks while offering trustworthiness for regulated industries.

Key concepts

Native Citation and Grounding
The models generate citations directly during inference using a Wikipedia-like tag syntax rather than relying on post-hoc methods. This provides higher control over how sources are presented and integrated into the final answer, ensuring evidence is intrinsically linked to the response.
Proto-agentic Reasoning Sequence
The models follow an iterative reasoning process covering trivial, standard, and refusal questions. This allows them to dynamically direct their own workflow based on query analysis and source evaluation, acting like a simple agent that decides the best path forward.
Tokenizer Recycling Method
To manage special tokens efficiently in smaller models, Pleias uses a new tokenizer variant. This method re-trains the last 19 tokens as special tokens to help the model quickly recognize instruction structures and mitigate performance drops associated with pre-allocated tokens.
Mid-training Data Generation
Models are mid-trained on synthetic RAG examples from the Common Corpus using a custom pipeline. This process scales training data, mitigates legal risks by reusing synthetic output, and ensures high quality through structured reasoning traces and adversarial examples.

Terminology used across episodes

This episode discusses

The paper

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family · Read on arXiv

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas Hinostroza, Matthieu Delsart, Irène Girard Othman Hicheur, Anastasia Stasenko, Ivan P. Yamshchikov

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Even Small Reasoners Should Quote Their Sources".

Jane: Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: The authors of "Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family" are introducing two small reasoning models, Pleias-RAG-350m and Pleias-RAG-1B <ref:2504.18225#pg0,Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model>. Their central thesis is that these models can outperform smaller models, even those under four billion parameters, on standardized RAG benchmarks like HotPotQA and 2WikiMultiHopQA <ref:2504.18225#pg0,4 billion parameters, on standardized RAG benchmarks>.

Jane: They claim these models offer native support for citation and grounding using literal quotes during the inference process, rather than relying on post-hoc citation methods that other works have focused on. This is a big difference in how they handle source integration.

Lu: What really stands out from their summary is that they've trained these mid-trained models on a large synthetic dataset emulating the retrieval of multilingual open sources from the Common Corpus, which totals about two trillion tokens.

Meng: That synthetic data generation approach sounds like it’s a smart way to tackle the data frictions we’ve seen with larger models, and I'm curious if that process helps them address those practical deployment issues on-device.

Lalam: I see the focus on systematically reference grounding across leading European languages as something that really elevates the capability of these smaller models beyond just performance metrics.

Tom: They also point out that these models are competitive with popular larger models like Qwen-two point five-7B, Llama-three point one-8B, and Gemma-three-4B while specifically maintaining consistent RAG performance across those European languages <ref:2504.18225#pg0,Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B>.

Jane: The paper highlights that because of their size and ease of deployment on constrained systems, these models are positioned as valuable for trustworthy AI applications in those specific settings.

Conclusion: Tom: Thinking about the title, "Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family," it really emphasizes that accuracy and sourcing shouldn't be reserved for massive models only, which is a key point for us to consider.

Jane: I think what this paper suggests is a design philosophy where the model’s output is inherently tied to verifiable evidence through direct quotes, making the entire system more auditable from the start.

Lu: The implication here, creatively speaking, is that we could design AI agents that are fundamentally more transparent in their reasoning because they aren't just guessing or synthesizing; they are citing what they found.

Meng: From a practical standpoint, if we can guarantee systematic reference grounding with literal quotes on these smaller models, it means we reduce the uncertainty when deploying them for tasks where reliability is paramount.

Lalam: The broader impact I see is that this approach could set a new standard for how we build AI systems intended for regulated industries because it directly addresses the verifiability gap and gives us evidence-based decision support.

Tom: So, essentially, the authors are showing that you don't need billions of parameters to achieve a solid foundation in RAG and source summarization if you prioritize this native citation structure.

Jane: Exactly, and they show that this approach works consistently across different European languages too, which opens up possibilities for more inclusive AI tools globally.

More episodes

← Home