PRISM: Parallel Residual Iterative Sequence Model
cs.LG
Submitted: 2026-02-11
Updated: 2026-09-15
Comments: 21 pages, 2 figures
Code: https://github.com/gpr-prism/prism
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models.
Terminology
Abstract
Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models. Existing efficient architectures are theoretically bounded by shallow, single-step linear updates, while powerful iterative methods like Test-Time Training (TTT) break hardware parallelism due to two dimensions of serial dependency: token-level state reliance and step-level iteration loops. We propose PRISM (Parallel Residual Iterative Sequence Model) to resolve this tension. PRISM explicitly approximates the expressive gate-residual-direction iteration pattern of TTT in a parallelizable form. We employ a Write-Forget Decoupling strategy that isolates non-linearity within the injection operator. To bypass the serial dependency of explicit solvers, PRISM utilizes a two-stage proxy architecture: a short-convolution anchors the initial residual using local history energy, while a learned predictor estimates the refinement updates directly from the input. This design distills structural patterns associated with iterative correction into a parallelizable feedforward operator. Theoretically, we prove that this formulation achieves Rank- L accumulation, structurally expanding the update scheme beyond the single-step Rank- 1 bottleneck. Empirically, it achieves comparable performance to explicit optimization methods while achieving 174x higher throughput. Codes are available in https://github.com/gpr-prism/prism/.
Sources
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Jamba: A Hybrid Transformer-Mamba Language Model
- RWKV-7 "Goose" with Expressive Dynamic State Evolution
- Titans: Learning to Memorize at Test Time
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Kimi Linear: An Expressive, Efficient Attention Architecture
- OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment
- Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences
- MoM: Linear Sequence Modeling with Mixture-of-Memories
- Gated Linear Attention Transformers with Hardware-Efficient Training
- Gated Delta Networks: Improving Mamba2 with Delta Rule
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
- Test-Time Training Done Right
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks