The Depth Flow of Token Representations Is Nonlinear and Does Not Descend Its Own Density
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- A Statistical Physics of Language Model Reasoning
- The Curved Spacetime of Transformer Architectures
- Transformer Dynamics: A neuroscientific approach to interpretability of large language models
- Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
- A mathematical perspective on Transformers
- The Remarkable Robustness of LLMs: Stages of Inference?
- A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases
- All-but-the-Top: Simple and Effective Postprocessing for Word Representations
- Sinkformers: Transformers with Doubly Stochastic Attention
- Lines of Thought in Large Language Models
- Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
- YuriiFormer: A Suite of Nesterov-Accelerated Transformers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering