Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
cs.AI, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 19 pages, 12 figures. Code and data: https://github.com/David-31415/maia-depth-migration. Built with chessformer-lens library: https://github.com/chessformer-lens/chessformer_lens
Code: https://github.com/David-31415/maia-depth-migration
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers.
Terminology
Abstract
Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.
Sources
- What Affects the Effective Depth of Large Language Models?
- Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
- Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
- Acquisition of Chess Knowledge in AlphaZero
- Aligning Superhuman AI with Human Behavior: Chess as a Model System
- Chessformer: A Unified Architecture for Chess Modeling
- Amortized Planning with Large-Scale Transformers: A Case Study on Chess
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Maia-2: A Unified Model for Human-AI Alignment in Chess
- Attention Is All You Need
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection