Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
cs.CL
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/OpenMOSS/Hybrid-Mechanics
Terminology
Sources
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen3 Technical Report
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
- gpt-oss-120b & gpt-oss-20b Model Card
- Gemma 4 Technical Report
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Kimi K3: Open Frontier Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- A Systematic Analysis of Hybrid Linear Attention
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Olmo Hybrid: From Theory to Practice and Back
- LLaMA: Open and Efficient Foundation Language Models
- SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
- Thus Spake Long-Context Large Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering