EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

arXiv:2603.02041 · cs.CL, cs.AI · Submitted 2026-03-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training".

Jane: The paper was written by Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu et al. from University of Tartu (Institute of Computer Science) and Tallinn University of Technology (Department of Software Science) and Institute of the Estonian Language and Tallinn University (School of Humanities).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: The paper outlines a very specific methodology in its summary, and it’s not just one step. They are taking Llama three point one 8B as their base model, which is a large, powerful LLM that already has some multilingual abilities.

Jane: But the core idea is that this base model needs more targeted exposure to Estonian, so they apply Continued Pretraining or CPT on a carefully constructed data mixture designed to teach it the language.

Lu: I’m fascinated by the composition of that mix—it's not just random Estonian text; it includes English replay and code and math, which is a very sophisticated way to ensure they don't lose its general knowledge.

Meng: From an engineering standpoint, that thirty-five point seven billion token mixture is quite impressive to manage; it shows how much data they needed to shift the model’s focus onto Estonian without just overwhelming it with sheer volume.

Lalam: It's about balancing that exposure—the blend of Estonian and general knowledge—to create a truly capable language model that respects the nuances of the language.

Tom: That mixture, containing eight point six billion tokens from the Estonian National Corpus, really forms the foundation of their work on "EstLLM."

Jane: It’s a very deliberate process, but we need to look at how this continued pretraining leads to actual changes in performance before we move onto the results.

Improvements: Tom: The paper's core findings demonstrate clear and consistent gains across various benchmarks for "EstLLM." They aren't just getting better at Estonian; they’re showing improvement in linguistic competence and knowledge too.

Jane: It seems the CPT process really builds a stronger foundation, but it also suggests that after continued pretraining, we need more specific tuning to make the model follow instructions well.

Lu: I was particularly interested in their use of "chat vector merging" to address this; it’s such an elegant way to inject the alignment properties from an instruction-tuned version back into the CPT model without having to retrain everything.

Meng: That technique, combining the base model with a chat vector difference, is very clever because it lets us recover instruction-following abilities while keeping that Estonian knowledge we just gained.

Lalam: The improvements in reasoning and translation are also impressive; they’re not just memorizing Estonian words, they’ are applying complex logic in the language.

Tom: And Jane mentioned that this isn's a perfect trade-off, right? The authors show that while EstLLM improves dramatically on Estonian tasks, it maintains competitive performance on English benchmarks.

Jane: Yes, but the fact that the results are consistent across automatic and comparative evaluations suggests a robust model without sacrificing its general capabilities.

Tom: That's a powerful finding; when we see improvements in one language without degradation in another area of achieving proficiency, it' something we should all be proud of "EstLLM."

Jane: It’s an encouraging result that leads us straight into the final wrap-up where we can discuss what this all means for global AI.

Conclusion: Tom: So, we have seen the journey of "EstLLM," from the initial problem of English dominance to how they solved it using CPT and post-training methods. The evidence is pretty compelling that these adaptation strategies work.

Jane: It's more than just a technical success for this specific project; it’ a roadmap for improving language equity in AI, which is a huge win for the whole team.

Lu: I think the implications are massive; if "EstLLM" proves successful, we are setting precedents for how diverse, smaller languages will be integrated into future global models.

Meng: Practically speaking, this demonstrates that we don're finding computationally inexpensive ways to achieve high performance after language adaptation without having to start from scratch every single time.

Lalam: This is about making sure the voices of all cultures are heard; "EstLLM" offers a path toward a more globally representative and culturally sensitive AI landscape for everyone.

Tom: I think we have a lot to process, but we want to wrap up this discussion on "EstLLM" by sharing our final thoughts with the listeners.

Jane: It's an exciting time for language technology, and it’s clear that this work is pushing the boundaries of what is possible.

Lu: I hope this research opens up possibilities for more diverse creative expressions in Estonian through AI too, which has immense potential for cultural growth.

Meng: I just hope we can see these methods applied to make a practical difference in deployment across different regions now.

Lalam: We need to keep the momentum going, ensuring that this success is the starting point, not the finish line for a broader language adaptation movement.

Tom: That's a great way to end our discussion on "EstLLM." Thank you all for helping us explore these concepts today!

Conclusion: Tom: So, we've covered so much ground today talking through "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training." It really shows how crucial localized data is for making these big models actually useful.

Jane: Exactly, Tom. What I take away from all of this is that it’s not enough just to have a massive model; you genuinely need to tailor it to the specific nuances of a language community, like Estonian.

Lu: That idea of tailoring really unlocks something incredible; we're talking about making these systems culturally aware, not just grammatically correct.

Meng: But practically speaking, Jane pointed out that local data is key—I wonder how much compute time it takes to successfully integrate a localized pretraining phase like this into an existing infrastructure?

Jane: Well, Meng, I think the breakthrough here isn't just the compute; it’s proving that continued pretraining *works* effectively enough to move the needle significantly.

Tom: And Lu was right about the cultural awareness—it's not just translation; it’s understanding context and local idioms, which is a huge leap for general AI models.

Lu: It suggests a paradigm shift where we view language models less as universal engines and more as deeply rooted digital extensions of specific cultures.

Meng: If that’s true, then the next hurdle for me is making the deployment scalable across dozens of small, specialized languages worldwide, not just one region.

Lalam: What Meng brings up about scaling is profound because it touches on how knowledge itself flows; if we can master fine-tuning for niche language groups, we elevate global human culture by giving voice to every dialect.

Jane: It makes you think about what happens when these tools become commonplace, doesn't it? They could genuinely democratize access to high-quality AI assistance everywhere.

Tom: Right, because the implications are huge—it means that smaller linguistic communities aren't left behind just because they don't speak English or Mandarin.

Lu: Looking forward, this work paves the way for hyper-localized digital assistants that can assist in everything from historical preservation to modern scientific fields within Estonia.

Meng: So, instead of just viewing it as a single model upgrade, we should see this methodology as a toolkit that any regional tech hub can adopt right away.

Lalam: Ultimately, the most impactful vision is seeing these enhanced models foster greater cross-cultural understanding by serving as bridges between language groups.

Jane: It's been such an exciting discussion wrapping up our look at "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training."

Tom: Thanks to all of you for breaking this down so well; we really appreciate the insights from Lu, Meng, and Lalam today.

Lu: It's been a wild ride through the possibilities!

Meng: Great discussion; I feel much clearer on the engineering path forward now.

Lalam: Remember that this advancement helps weave a richer tapestry of global human expression.

Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fišel, Tanel Alumäe, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts

University of Tartu · Tallinn University of Technology · Institute of the Estonian Language · Tallinn University

cs.CL, cs.AI

Submitted: 2026-03-02

Updated: 2026-08-25

Code: https://github.com/tlu-dt-nlp/EstGEC-L2-Cor

Importance score: 88/100

The gist: The paper addresses the critical need to enhance the capabilities of large multilingual language models (LLMs) specifically for Estonian, a language that often lacks sufficient digital resources

Key concepts

EstLLM
EstLLM is a large language model designed to enhance Estonian capabilities. It is built upon the Llama 3.1 8B base model, which was then adapted using Continued Pretraining (CPT) on a curated dataset to make the model highly proficient in Estonian.
Continued Pretraining (CPT)
This is a methodology used in the paper where a base LLM is given targeted exposure to Estonian language data. It involves training the model on specific, high-quality text to teach it nuances and improve its linguistic competence without losing general knowledge.
Chat Vector Merging
This technique was used after CPT to address instruction-following issues. It allows researchers to inject alignment properties from an instruction-tuned version back into the CPT model, helping the model follow commands while retaining its new Estonian knowledge.

Terminology

Summary

The paper addresses the critical need to enhance the capabilities of large multilingual language models (LLMs) specifically for Estonian, a language that often lacks sufficient digital resources compared to major global languages. The study proposes methods involving continued pretraining and model merging using extensive Estonian corpora, aiming to significantly boost performance on local benchmarks and improve overall utility in real-world applications.

Continued Pretraining and Model Merging Techniques

The research evaluates the effectiveness of adapting existing foundational models through specialized training methods. One primary approach involves continued pretraining with Estonian corpora (monolingual), which is demonstrated across various discriminative benchmarks. Furthermore, the authors investigate model combination via SLERP merging, where the resulting model is a blend between two base models: Llama-3.1-8B (t=0) and Llama-3.1-EstLLM-8B (t=1). This merging technique allows for systematic testing of how different degrees of Estonian specialization influence overall performance metrics.

Performance on Structured Benchmarks

The efficacy of the EstLLM enhancements is rigorously tested using a comprehensive suite of Estonian discriminative benchmarks, including:

  • Belebele ET and Exam ET (general language understanding).

  • Inflection ET, Trivia ET, Winogrande ET, and XCOPA ET (testing specific knowledge domains).

  • Grammar ET and GlobalPIQA ET (assessing grammatical accuracy and general knowledge).

For broader linguistic capabilities, the model performance is also measured on established multilingual benchmarks such as MMLU-Redux. The evaluation of cross-lingual transfer is demonstrated using FLORES200 EN-ET and FLORES200 ET-EN, providing quantitative metrics for both English Estonian translation and understanding.

Real-World Open Model Ranking in Estonian

Beyond controlled benchmarks, the paper provides a snapshot of open model performance on the AI Barometer—a Chatbot Arena style evaluation page in Estonian—as of 19.02.2026 (Table 13). This leaderboard ranks various open models based on their practical capabilities in the target language. The models evaluated include DeepSeek-V3, Kimi-K2-Instruct, Llama 4 Maverick, and multiple iterations of EstLLM.

The ranking provides a detailed view of model performance across several metrics:

  • Model size: Ranging from 7B to 405B parameters.

  • Score: Indicating measured capability (e.g., DeepSeek-V3 scores 1457).

  • 95% CI: Representing the confidence interval of the score.

The data reveals that models specifically trained or enhanced for Estonian, such as those labeled Llama-EstLLM, are positioned within a competitive landscape alongside massive generalist models (e.g., Meta-Llama-3.1-405B-Instruct), demonstrating the measurable impact of targeted domain adaptation on model utility in Estonian.

Improvements for AI systems

Improvement: Instead of relying solely on Spherical Linear Interpolation (SLERP) for merging two models (M Base and M EstLLM), we must develop a Task-Weighted Merging (TWM) framework. This framework modifies the standard merging process by incorporating task-specific gradients during the interpolation phase.

Mechanism:

  1. Identify key, orthogonal linguistic domains relevant to Estonian (e.g., morphology/declension, syntax/grammar, specialized vocabulary/trivia).

  2. During the SLERP calculation of weights W(t), we calculate not just a linear blend based on the magnitude of t, but a weighted gradient grad L task(t).

  3. The merging objective function is altered to: W (L General(M) + lambda 1 L Grammar(M) + lambda 2 L Trivia(M) +...), where the loss terms L are sampled from the respective Estonian benchmarks (Table 12).

  4. This ensures that the merged model is not just an average of two models, but a specialized synthesis optimized for high performance across all critical linguistic axes simultaneously.

Improved System Capability: The resulting AI system will exhibit Domain-Specific Resilience. It will maintain state-of-the-art performance in niche, difficult tasks (like Belebele Grammar Declension or Exam Belebele) that are often the failure points of generalist models, while retaining the general knowledge base of the larger foundational model.


Sources

Related papers