RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
cs.LG, cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/acetocarmine11/rope-profiler
Terminology
Sources
- Extending Context Window of Large Language Models via Positional Interpolation
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
- The Llama 3 Herd of Models
- Gemma 3 Technical Report
- Disentangling the Expressivity of RoPE
- Kimi K3: Open Frontier Intelligence
- CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
- LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
- Ministral 3
- Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
- Untwisting RoPE: Frequency Control for Shared Attention in DiTs
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design
- Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
- DoPE: Denoising Rotary Position Embedding
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks