What Makes Good Multilingual Reasoning? Disentangling Traces with Measurable Features
cs.CL, cs.AI
Submitted: 2026-04-06
Updated: 2026-09-20
Comments: COLM 2026
Code: https://github.com/dayeonki/multilingual_reasoning
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Reasoning Models (LRMs) still exhibit large performance gaps between English and other languages, yet much current work assumes these gaps can be closed simply by making reasoning in every
Terminology
Abstract
Large Reasoning Models (LRMs) still exhibit large performance gaps between English and other languages, yet much current work assumes these gaps can be closed simply by making reasoning in every language resemble English reasoning. This work challenges this assumption by asking instead: what actually characterizes successful reasoning traces in multilingual settings, and to what extent do English-derived reasoning features genuinely help in other languages? We first define a suite of measurable reasoning features spanning multilingual alignment, reasoning step, and reasoning flow aspects of reasoning traces, and use logistic regression to quantify how each feature associates with final answer accuracy. We further train sparse autoencoders over multilingual traces to automatically discover latent reasoning concepts that instantiate or extend these features. Finally, we use the features to re-rank traces and measure their impact on accuracy at test time. Across two mathematical reasoning benchmarks, four LRMs, and ten languages, we find that most features are positively associated with accuracy, but the strength of association varies considerably across languages and can even reverse in some. Our findings challenge English-centric reward designs and point toward adaptive objectives that accommodate language-specific reasoning patterns, with concrete implications for multilingual benchmark and reward design.
Sources
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Multilingual Reasoning Gym: Multilingual Scaling of Procedural Reasoning Environments
- Aligning Multilingual Reasoning with Verifiable Semantics from a High-Resource Expert Model
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Could Thinking Multilingually Empower LLM Reasoning?
- ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
- TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning
- Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
- Measuring Faithfulness in Chain-of-Thought Reasoning
- LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
- THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
- R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning
- Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering