The Geometry of Tokens in Internal Representations of Large Language Models
Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga, Matteo Biagetti
cs.CL, cs.LG
Submitted: 2026-08-21
Updated: 2026-08-24
Comments: 12+14 pages, 18 figures, matches published version on Transactions of Machine Learning Research
Journal ref: Transactions of Machine Learning Research, ISSN: 2835-8856, 2026, https://openreview.net/pdf?id=rBEgNAslpY
Code: https://github.com/RitAreaSciencePark/token
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Mathematical Theory of Attention
- A mathematical perspective on Transformers
- Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
- Generic controllability of equivariant systems and applications to particle systems and neural networks
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Emergence of a High-Dimensional Abstraction Phase in Language Transformers
- Persistent Topological Features in Large Language Models
- Levels of Analysis for Machine Learning
- Measure-to-measure interpolation using Transformers
- Dynamic metastability in the self-attention model
- Unsupervised detection of semantic correlations in big data
- Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
- Distributional Results for Model-Based Intrinsic Dimension Estimators
- Mistral 7B
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Loss Landscape Degeneracy and Stagewise Development in Transformers
- Lines of Thought in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering