DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

arXiv:2608.13517 · cs.CL, cs.AI · Submitted 2026-08-20 · Read on arXiv

Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech

University of Southern Denmark · Ordbogen A/S · Aarhus University

cs.CL, cs.AI

Submitted: 2026-08-20

Updated: 2026-08-24

Comments: Technical Report, 20 Pages, 1 Model, Hierarchical Reasoning Model

Code: https://github.com/schneiderkamplab/HRM-Text

Project page: https://odensenlp.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: DFM Mimir v1 is a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, trained from scratch and delivering "highly competitive performance for English and

Terminology

Summary

DFM Mimir v1 is a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, trained from scratch and delivering highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. The model is trained on a mixture of 161 datasets totaling approximately 70.5 billion tokens per epoch, and it "outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish."

The work addresses the challenge that current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. The authors employ the HRM-Text framework to enable focusing on post-training data during the initial training phase, thereby facilitating the creation of a viable base model for the wider community. Since some datasets used in the original HRM-Text work do not align with DFM's permissibility standards, they demonstrate the efficacy of synthetically generating 'transplant datasets', replacing non-permissible data with synthetically generated permissible variants, finding that these synthetic alternatives achieve comparable or superior performance without compromising data rights.

The training corpus is curated into eight functional categories: Danish instruction & knowledge (22.07% of tokens), English instruction (19.26%), Sapient mixed (17.02%), Math & reasoning (14.76%), Mimir synthetic (10.00%), Agentic & tool use (9.46%), Machine translation (4.96%), and Science & summarization (2.47%). The corpus is predominantly English (68.62%), with Danish contributing 24.74% and bilingual Danish–English data a further 6.54%. Data enters the corpus in seven forms: reformatted (65.96%), curated + reformatted (16.91%), synthetic + audited (11.08%), tool-call formatted (2.65%), translated + audited (2.26%), agreement-supplied (0.95%), and derived task (0.18%). The corpus is highly concentrated, with the top ten datasets accounting for 66.5% of all tokens, and the top three alone for 38.1%. Some datasets are repeated multiple times per epoch, such as lærebogen (4×) and eight small Danish datasets (10×).

The architecture uses a hidden size of 1,536, 32 layers, 12 attention heads per layer, a feed-forward expansion factor of 4, 2 H-cycles and 3 L-cycles, truncated backpropagation limited to 5 steps with a warmup ratio of 0.2, Rotary Position Embedding (RoPE) with θ = 10,000, and pre-norm layer normalisation with ϵ = 10−6. The model is trained from scratch using the Gemma-4 tokenizer, with Fully Sharded Data Parallelism (FSDP) using bfloat16 for computation and fp32 for gathering precision. Training uses the AdamW optimizer with a peak learning rate of 3 × 10−4, 2,000-step linear warm-up, constant schedule thereafter, global batch size of 262,144 tokens, gradient accumulation of 2 on 8 accelerators, and was run for 1.65M steps on 8 NVIDIA B200 GPUs in just under 3 weeks.

On English benchmarks, Mimir outperforms all considered competitors on BoolQ, Winogrande, and DROP. On Math & Code, Mimir leads across its weight-class for GSM8K and HumanEval, with Mimir being second overall on GSM8K and better than Qwen3.5 2B on HumanEval. On Danish benchmarks, Mimir outperforms all competitors on grammatical tasks (DaLA, GEC), question-answering tasks (WikiQA), and is close to the best on Nordjylland News (N.News; summarization). On average, "Mimir displays superior performance on the Danish benchmarks, is only 0.3 points behind Qwen 3.5 4B on English tasks, and only 3.8% behind SmolLM3 3B on Math & Code." On Math & Code, Mimir yields a 36.7% improvement compared to HRM-Text (64.1 Mimir vs. 46.9 HRM-Text).

The conclusion states that Mimir v1 "provides frontier-level performance using only permissible data, excluding data containing personal information or copyright infringement and including data that is either openly licensed, made available by agreement, or allowed by the European Union's text and data mining exception for research institutions. The authors note that Mimir v1 still lags behind Gemma 4 (5B, effective 2.3B) on the Math & Code domains, making room for improvement in future iterations. Future work will focus on investigating scaling behavior of HRM models like Mimir, exploring reinforcement learning (which is yet unexplored for this architecture), and continuing to develop the dataset to achieve full openness regarding licensing and further improved model performance."

Improvements for AI systems

Improvements to AI Systems:

  1. Permissible-Data Training Pipeline
  • Implement a synthetic transplant dataset generator that replaces non-permissible training data with ethically sourced, synthetically generated equivalents, achieving comparable or superior performance without copyright or privacy violations.

  • Enable automated auditing of training corpora against permissibility standards (e.g., EU TDM exception, open licenses) to ensure legal compliance by default.

  1. Hierarchical Reasoning Model (HRM) with Post-Training Focus
  • Adopt the HRM architecture (2 H-cycles, 3 L-cycles) to prioritize post-training data during initial training, reducing reliance on massive pre-training corpora and lowering the barrier for open-source research.

  • Integrate truncated backpropagation (5 steps) with a warmup ratio of 0.2 to stabilize training on smaller, curated datasets, improving sample efficiency.

  1. Efficient Scaling on Limited Hardware
  • Use the training recipe (FSDP with bfloat16/fp32, AdamW at 3e-4 peak LR, global batch 262K tokens) to train a 1B-parameter model on 8 GPUs in under 3 weeks, enabling reproducible training for institutions with modest compute.

  • Apply the Gemma-4 tokenizer and RoPE (θ=10,000) to achieve strong multilingual performance (Danish + English) without additional tokenizer training.

  1. Category-Aware Data Curation
  • Structure training data into eight functional categories (e.g., Danish instruction, agentic/tool use, math & reasoning) with explicit token budgets, allowing dynamic re-weighting to target weak domains (e.g., increase Math & Code share to close the 3.8% gap with SmolLM3 3B).

  • Implement repeated dataset sampling (e.g., 4× for lærebogen, 10× for small Danish sets) to reinforce low-resource language and task coverage.

  1. Multilingual Benchmark Optimization
  • Leverage Mimir’s superior Danish performance (DaLA, GEC, WikiQA) to build a Danish-first assistant that outperforms larger models (Qwen 3.5 4B) on grammatical and QA tasks, while maintaining competitive English results (within 0.3 points of Qwen 3.5 4B).

  • Use the 36.7% Math & Code improvement over HRM-Text to create a code-generation and reasoning model that rivals 2B–4B competitors (e.g., better than Qwen3.5 2B on HumanEval).

What the Improved AI System Can Do:

  • Legally Deployable Assistant: Generate responses, code, and summaries using only ethically sourced data, making it safe for commercial use in the EU and beyond without copyright or privacy litigation risks.

  • Low-Resource Language Specialist: Provide state-of-the-art Danish grammar correction, question answering, and summarization, outperforming models 4× its size, while maintaining solid English and code capabilities.

  • Resource-Efficient Research Platform: Enable small teams to train competitive 1B models from scratch in 3 weeks on 8 GPUs, democratizing LLM development for academia and startups.

  • Adaptive Task Router: Dynamically allocate training tokens across categories (e.g., boost agentic/tool use for API-calling tasks, or math for reasoning-heavy workloads) based on downstream benchmark feedback.

  • Synthetic Data Augmenter: Automatically generate permissible replacements for any non-compliant dataset, allowing existing corpora to be reused legally while preserving or improving model quality.

  • Benchmark-Tuned Reasoner: Achieve top scores on BoolQ, Winogrande, DROP, GSM8K, and HumanEval within its weight class, making it suitable for reasoning, QA, and code generation in production settings.

Sources

Related papers