Disentangling the Expressivity of RoPE
Selim Jerad, Anej Svete, Jiaoda Li, Ryan Cotterell
Toyota Technological Institute at Chicago · ETH Zürich
cs.LG, cs.FL
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: This paper, "Disentangling the Expressivity of RoPE" by Selim Jerad, Anej Svete, Jiaoda Li, and Ryan Cotterell, published at COLM 2026, reconciles two seemingly incompatible accounts of rotary
Terminology
Summary
This paper, Disentangling the Expressivity of RoPE
by Selim Jerad, Anej Svete, Jiaoda Li, and Ryan Cotterell, published at COLM 2026, reconciles two seemingly incompatible accounts of rotary position embeddings (RoPE) in transformers: the theoretical expressivity view linking periodic position information to modular predicates, and the mechanistic/long-context view emphasizing positional anchors and local offsets.
The authors formalize both accounts for fully uniform, finite-precision soft-attention transformers. Their central finding is that if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates
(LTL[P, MOD] = PFO2[<, MOD]). In contrast, Conventional RoPE is different: The rotations it computes never repeat,
yielding a precision-dependent bounded simulation of fixed-offset lookback operators, rather than an all-length modular characterization.
Key theoretical contributions:
-
Characterization of periodic RoPE (RoPEP): The paper proves Theorem 3.1: SMAT[RoPEP] = LTL[P, MOD] = PFO2[<, MOD]. This means component-periodic RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates. The proof shows that periodic schedules can be expressed via finitely many modular predicates (Proposition 2.1), and conversely, each modular atom can be implemented using a period-m rotary pair with a lookup table indexed by position modulo the period.
-
Equivalence with periodic sinusoidal encodings: Theorem 3.2 shows SMAT[SiPEP] = LTL[P, MOD], demonstrating that
the key factor for the corresponding all-length logical class is the periodicity of the encoding; both the absolute and relative encodings are unified under modular predicates.
-
Rich characterization of LTL[P, MOD]: Theorem 3.3 provides five equivalent characterizations: (i) left-deterministic modular polynomials, (ii) syntactic monoids in QR, (iii) existence of k such that the k-automaton is partially ordered, (iv) PFO2[<, MOD] definability, and (v) LTL[P, MOD] definability.
-
Inexpressibility results: The paper proves that LTL[P, MOD] and LTL[S] are incomparable. For instance, (aa)* is in LTL[P, MOD] but not star-free, while ΣabΣ (with Σ > 2) is locally testable but not in LTL[P, MOD]. PARITY (even number of 1s) is shown to be outside LTL[P, MOD] since
modular predicates count positions, not symbols.
-
Non-periodic RoPE analysis: Proposition 4.1 shows that irrational rotations imply non-periodicity. The paper then proves (Corollary 4.1) that with irrational rotations, RoPENP transformers can implement fixed-offset attention heads up to a certified length bound, providing
a bounded simulation of LTL[P, Y]
rather than an all-length characterization.
Experimental findings:
The authors train transformer classifiers on controlled formal languages (training on strings up to length 40, testing on lengths 41–500):
-
Periodic constructions: RoPEP achieves perfect generalization (accuracy 1.00, N* = 500) on both (aa)* and (ab)*, matching the modular-predicate construction. SiPEP struggles on these tasks despite having the same existential expressivity, suggesting
absolute sinusoidal features in the residual stream are more easily overwritten.
-
Conventional non-periodic rotations: Neither RoPENP nor SiPENP generalizes perfectly on the two periodic languages across a sweep of bases 10−12, 10−8, 10−4, 1, 104, 108, 1012,
separating the engineered periodic behavior from conventional frequencies.
-
Locality account: On the LTL[Y] language Σa, RoPENP's mean N exceeds NoPE's at six of seven bases, supporting the fixed-offset mechanism. However, on the LTL[P] language aΣ*,
the NoPE mean is 470.5, whereas the RoPENP means range from 146.9 to 402.9,
and on ΣabΣ,the NoPE mean is 92.9, whereas the RoPENP means range from 46.1 to 61.9.
This suggestsa tradeoff: Relative displacement can help retrieve a nearby symbol while disrupting the nearly position-invariant aggregation that supports conditioning on distant information.
-
Base sensitivity: The standard base β = 104 gives the lowest mean N* among tested bases on single-operator languages (45.5 on Σa and 146.9 on aΣ), mirroring prior extrapolation experiments.
Connection to attention sinks: The construction for modular predicates uses BOS as a fixed anchor, connecting to attention sinks. The paper notes Gemma 7B, which uses RoPE, does concentrate attention on BOS,
suggesting sinks in RoPE transformers serve in part to expose positional phase.
Limitations and implications: The periodic lower bound uses a repeated finite table and tailored exact value set, not implying standard frequencies discover modular behavior. The locality result is bounded, not an all-length characterization. The paper suggests reserving unrotated dimensions when a model must combine global P reasoning with periodic or local relative-position information
and conjectures that partial-RoPE and NoPE-RoPE hybrid attention variants should improve long-context P reasoning while retaining locality bias.
Improvements for AI systems
Improvement 1: Adaptive Positional-Encoding Scheduler for Long-Context Transformers
-
What to change: Replace fixed RoPE frequencies with a hybrid scheduler that dynamically allocates a subset of rotary dimensions to periodic (integer-multiple) frequencies and the rest to non-periodic (irrational) frequencies, based on the task’s detected need for modular vs. local reasoning.
-
Improved capability: The system can simultaneously handle (a) exact modular counting (e.g., detecting
(aa)*or(ab)*patterns at any length) and (b) fine-grained local offset retrieval (e.g., finding the previous symbol) without the tradeoff observed in the paper. It will generalize perfectly on both periodic languages and local-context tasks, unlike current RoPE which fails on one or the other.
Improvement 2: Attention-Sink-Aware Positional Phase Exposure
-
What to change: Modify the attention mechanism to explicitly expose the positional phase of the BOS token (as an attention sink) by adding a learnable phase-vector projection that is always attended to, and by initializing the BOS embedding to contain a full period’s worth of phase information.
-
Improved capability: The system will use BOS as a reliable anchor for modular predicates, enabling exact long-range periodic pattern recognition (e.g.,
(ab) nfor n > 500) without requiring the model to learn this from scratch. This reduces the need for massive training data on long sequences and improves few-shot generalization to unseen lengths.
Improvement 3: Task-Aware Frequency Base Selection
-
What to change: Implement a meta-learning module that, given a task description or a small validation set, selects the optimal RoPE base β (from a continuous range) by estimating whether the task requires modular (periodic) or local (offset) reasoning. For modular tasks, use β that makes rotations repeat (e.g., β = 1); for local tasks, use β = 104 or higher.
-
Improved capability: The system will automatically adapt its positional encoding to the task at hand, avoiding the observed performance collapse (e.g., NoPE outperforming RoPE on
aΣ*andΣ*abΣ*). It will match or exceed the best fixed-base performance on both periodic and local-context benchmarks, with no manual tuning.
Improvement 4: Hybrid NoPE-RoPE Attention Heads for Long-Context P Reasoning
-
What to change: In each transformer layer, dedicate a fixed fraction of attention heads to NoPE (no positional encoding) and the rest to RoPE, with a learned gating mechanism that routes tokens to the appropriate head type based on whether the current context requires position-invariant aggregation (e.g., counting symbols) or position-sensitive retrieval.
-
Improved capability: The system will retain RoPE’s locality bias for nearby token interactions while gaining NoPE’s ability to aggregate distant information without disruption. This directly addresses the paper’s finding that RoPE disrupts position-invariant aggregation, enabling superior performance on tasks like
aΣ*(where NoPE excels) andΣ*abΣ*(where RoPE fails) while still handling long-context local retrieval.
Improvement 5: Precision-Adaptive Rotation for Bounded Lookback
-
What to change: For non-periodic RoPE, instead of using a fixed precision, dynamically adjust the numerical precision of the rotary phase based on the current sequence length, ensuring that the certified bound on fixed-offset lookback (from Corollary 4.1) is maintained even for very long sequences.
-
Improved capability: The system will reliably simulate fixed-offset attention (e.g.,
Yoperator) up to a length that scales with available precision, avoiding the degradation seen in standard RoPE at long contexts. This is particularly useful for streaming or memory-constrained applications where sequence length is unbounded.
Improvement 6: Modular-Predicate-Aware Training Objective
-
What to change: Add an auxiliary loss term during training that encourages the model’s positional phase to align with modular predicates (e.g., by predicting
position mod mfor a set of small m values) for a subset of rotary dimensions, while the main task loss handles the primary objective. -
Improved capability: The system will learn to use periodic positional information more efficiently, achieving perfect generalization on modular languages with fewer training examples and shorter training sequences (e.g., training on lengths ≤ 20 instead of 40), as the auxiliary objective pre-structures the phase space.
Improvement 7: Runtime Detection of Periodicity vs. Locality Needs
-
What to change: Implement a lightweight classifier at inference time that analyzes the input sequence’s statistical properties (e.g., autocorrelation, symbol repetition frequency) to decide whether to engage periodic RoPE (for modular tasks) or non-periodic RoPE (for local tasks), and switch the positional encoding accordingly.
-
Improved capability: The system will dynamically adapt to mixed-task inputs (e.g., a document containing both periodic code patterns and free-text local dependencies), maintaining high accuracy on both without retraining. This enables deployment in heterogeneous real-world scenarios where task type is unknown a priori.
Sources
- Qwen Technical Report
- Round and Round We Go! What makes Rotary Positional Encodings useful?
- Why do LLMs attend to the first token?
- Tighter Bounds on the Expressivity of Transformer Encoders
- Adding modular predicates to first-order fragments
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Neural Networks and the Chomsky Hierarchy
- RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
- The Llama 3 Herd of Models
- The Impact of Positional Encoding on Length Generalization in Transformers
- Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
- Characterizing the Expressivity of Fixed-Precision Transformer Language Models
- Characterizing the Expressivity of Local Attention in Transformers
- Exposing Attention Glitches with Flip-Flop Language Modeling
- Scaling Laws of RoPE-based Extrapolation
- Base of RoPE Bounds Context Length
- Why Are Linear RNNs More Parallelizable?
- Olmo Hybrid: From Theory to Practice and Back
- Rational Transductors
- Olmo 3
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks