SketchSSM: Write to the Full State, Read from a Compact Sketch
cs.LG, cs.DC
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/Johnny-Liou/ReplaySSM
Terminology
Sources
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Jamba: A Hybrid Transformer-Mamba Language Model
- Let's Verify Step by Step
- The Key to State Reduction in Linear Attention: A Rank-based Perspective
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
- DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
- KVBuffer: IO-aware Serving for Linear Attention
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks