Low-Bit Recurrent States in Hybrid Language Models
cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
- Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
- KVBuffer: IO-aware Serving for Linear Attention
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
- OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond
- Microscaling Data Formats for Deep Learning
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
- Pointer Sentinel Mixture Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks