RoPE attention is an exact forward-pass gradient step with softmax intact
stat.ML, cs.LG
Submitted: 2026-09-06
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: We derive an exact gradient-step representation of the RoPE-softmax forward pass.
Terminology
Abstract
We derive an exact gradient-step representation of the RoPE-softmax forward pass. For every deterministic RoPE-softmax attention head with arbitrary affine projection weights, we construct a query-dependent effective matrix ΔM i satisfying y i = μ i + u i ΔM i, where μ i is the uniform mean of the attended values and u i is the augmented query input. The construction applies the classical exponential divided difference ρ= ϕ 1 to retain the softmax exactly. Its positive coefficients give a unit gradient-step representation on a query-conditioned quadratic objective. The same function connects the RoPE generator to exact positional finite differences. We derive a tokenwise formula for the error of reusing one query's matrix and prove that a nonconstant finite-cache head cannot admit a globally exact affine query readout. Reconstruction checks and frozen-reuse calibration on one pretrained Qwen2.5-0.5B layer verify the representation and quantify the correction required when one query's matrix is reused.
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey