Positional Encoding via Token-Aware Phase Attention
cs.CL, cs.AI
Submitted: 2025-09-16
Updated: 2026-09-06
Comments: 28 pages
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scores that limits RoPE's ability to model long-context.
Terminology
Abstract
We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scores that limits RoPE's ability to model long-context. RoPE extension methods may alleviate this issue, but they typically require post-hoc adjustments after pretraining, such as rescaling or hyperparameters retuning. This paper introduces Token-Aware Phase Attention (TAPA), a new positional encoding method that incorporates a learnable phase function into the attention mechanism. TAPA preserves token interactions over long range, extends to longer contexts with direct and light continual pretraining, extrapolates to unseen lengths, and attains substantially lower perplexity and stronger retrieval performance in the long-context regime than RoPE-style baselines.
Sources
- Phi-4 Technical Report
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Extending Context Window of Large Language Models via Positional Interpolation
- PaLM: Scaling Language Modeling with Pathways
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- LongNet: Scaling Transformers to 1,000,000,000 Tokens
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Transformer Language Models without Positional Encodings Still Learn Positional Information
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Mistral 7B
- The Impact of Positional Encoding on Length Generalization in Transformers
- FNet: Mixing Tokens with Fourier Transforms
- Scaling Laws of RoPE-based Extrapolation
- Decoupled Weight Decay Regularization
- Mesa-Extrapolation: A Weave Position Encoding Method for Enhanced Extrapolation in LLMs
- RWKV: Reinventing RNNs for the Transformer Era
- YaRN: Efficient Context Window Extension of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering