A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

arXiv:2608.29034 · cs.CL, cs.AI · Submitted 2026-08-29 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-08-29

Updated: 2026-08-29

Comments: 32 pages, 6 figures

License: http://creativecommons.org/licenses/by/4.0/

The gist: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings.

Terminology

Abstract

A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure could be represented in vector space --- as filler-role bindings. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models --- from small toy models to LLMs --- to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.

Related papers