AdaMame: A Training Recipe for Adaptive Multilingual Reasoning
cs.CL, cs.AI
Submitted: 2026-06-13
Updated: 2026-09-20
Comments: EMNLP 2026
Code: https://github.com/dayeonki/adamame
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomenon known as language collapse.
Terminology
Abstract
While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomenon known as language collapse. Existing RL-based fixes typically add a binary language fidelity reward to the accuracy objective, yet still incur trade-off in accuracy, mid-trace code-switching, and excessive token usage. In this work, we propose AdaMame, a two-stage training recipe for multilingual mathematical reasoning that addresses these limitations by adaptively aligning the reasoning language to the query language without compromising accuracy. The first SFT stage fine-tunes on non-MT reasoning traces across five languages to establish multilingual reasoning capability. In the subsequent RL stage, we introduce AdaMame-GRPO, an adaptation of Group Relative Policy Optimization (GRPO) in which a query-conditioned alignment factor grows progressively during training, guiding the model to first explore diverse reasoning languages before exploiting reasoning in the query language. Evaluated across two benchmarks, two LRMs, and 12 languages, AdaMame-GRPO achieves Pareto-optimal performance across reasoning accuracy, language fidelity, and token efficiency over all baselines, with the strongest gains on out-of-domain, lower-resource languages.
Sources
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Could Thinking Multilingually Empower LLM Reasoning?
- ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
- ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
- LoRA: Low-Rank Adaptation of Large Language Models
- Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
- What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?
- What Makes Good Multilingual Reasoning? Disentangling Traces with Measurable Features
- R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning
- Understanding R1-Zero-Like Training: A Critical Perspective
- Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
- Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- Qwen2.5 Technical Report
- Magistral
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Language Models are Multilingual Chain-of-Thought Reasoners
- OpenAI GPT-5 System Card
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering