Rethinking the Multilingual Reasoning Gap with Layer Swap
cs.CL
Submitted: 2026-05-26
Updated: 2026-08-26
Code: https://github.com/huggingface/trl
Terminology
Sources
- Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models
- Long Chain-of-Thought Reasoning Across Languages
- Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
- Olmo 3
- A Survey of Multilingual Reasoning in Language Models
- Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
- Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
- ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
- Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Do Multilingual LLMs Think In English?
- OpenAI o1 System Card
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Language Models are Multilingual Chain-of-Thought Reasoners
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering