Multilinguality in Hybrid Attention LLMs
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/ibm-granite/granite-4.0-language-models
Terminology
Sources
- What Attention Recalls and Recurrence Controls in Hybrid Language Models
- Neural Machine Translation by Jointly Learning to Align and Translate
- Longformer: The Long-Document Transformer
- OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- Language-Specific Latent Process Hinders Cross-Lingual Performance
- Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Olmo 3
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Rethinking the Role of Efficient Attention in Hybrid Architectures
- Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
- Tiny Aya: Bridging Scale and Multilingual Depth
- Kimi Linear: An Expressive, Efficient Attention Architecture
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering