Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accounts
cs.CL
Submitted: 2026-05-20
Updated: 2026-09-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: Brain-language model alignment is often interpreted as evidence that transformer models implement computations similar to those of the human brain.
Terminology
Abstract
Brain-language model alignment is often interpreted as evidence that transformer models implement computations similar to those of the human brain. This assumes that neural predictivity reflects internal computational properties of large language models (LLMs), such as hierarchical contextual processing, predictive coding, or representational compression. An alternative possibility is that brain scores primarily reflect stable lexical-semantic correspondences shared by language models and the brain. Here we tested these interpretations using whole-brain encoding models across Mandarin, English, and French. Across all three languages, transformer representations significantly predicted activity in a distributed network spanning classical language regions, transmodal cortical systems, and subcortical structures. These spatial patterns showed substantial cross-linguistic overlap and remained remarkably stable across layers, providing little evidence that model depth systematically maps onto cortical processing hierarchies. Likewise, contextual transformer embeddings did not consistently outperform static lexical embeddings, despite providing some unique predictive variance. Finally, neither surprisal nor intrinsic dimensionality reproduced the layer-wise profile of brain scores, arguing against prediction and information compression as primary explanations for brain-LLM alignment. Together, these findings suggest that brain-LLM alignment is more robust across languages, transformer depth, and model architectures than previously appreciated, but less informative about shared computational mechanisms. Our results are more consistent with neural predictivity reflecting stable representational structure preserved across model transformations than with a one-to-one correspondence between their underlying computations.
Sources
- Low-Dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
- Scaling laws for language encoding models in fMRI
- Contextual Embeddings: When Are They Worth It?
- Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations
- Evidence from fMRI Supports a Two-Phase Abstraction Process in Language Models
- Emergence of a High-Dimensional Abstraction Phase in Language Transformers
- Abstraction Induces the Brain Alignment of Language and Speech Models
- What Are Large Language Models Mapping to in the Brain? A Case Against Over-Reliance on Brain Scores
- The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
- Action and Perception as Divergence Minimization
- Do Large Language Models Think Like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI
- How is BERT surprised? Layerwise detection of linguistic anomalies
- Joint processing of linguistic properties in brains and language models
- Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps
- Scaling and context steer LLMs along the same computational path as the human brain
- A Primer in BERTology: What we know about how BERT works
- Layer by Layer: Uncovering Hidden Representations in Language Models
- Attention Is All You Need
- Exploring Similarity between Neural and LLM Trajectories in Language Processing
- Brains and language models converge on a shared conceptual space across different languages
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering