Thomson: Continual Learning of Frontier Models for SovereignAI
cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Open-weight model: https://huggingface.co/thomsonreuters/Thomson-1.0-Small
Code: https://github.com/bytedance/deer-flow
Project page: https://langchain-ai.github.io/langgraph
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between
Terminology
Abstract
The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive π-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
- General Scales Unlock AI Evaluation with Explanatory and Predictive Power
- Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Deep Research Agents: A Systematic Examination And Roadmap
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- ReAct: Synergizing Reasoning and Acting in Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Refusal in Language Models Is Mediated by a Single Direction
- Balancing Continuous Pre-Training and Instruction Fine-Tuning: Optimizing Instruction-Following in LLMs
- Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
- Functional Regularisation for Continual Learning with Gaussian Processes
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- Qwen3 Technical Report
- The Llama 3 Herd of Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection