Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation

arXiv:2408.04619 · cs.LG, cs.AI, cs.CL, cs.HC · Submitted 2026-08-10 · Read on arXiv

Aeree Cho, Grace C. Kim, Alexander Karpekov, Seongmin Lee, Alec Helbling, Benjamin Hoover, Zijie J. Wang, Minsuk Kahng, Duen Horng (Polo) Chau

Georgia Institute of Technology · IBM Research AI · Yonsei University

cs.LG, cs.AI, cs.CL, cs.HC

Submitted: 2026-08-10

Updated: 2026-08-11

Comments: CHI 2026 full paper. Extended version of the 2-page paper presented at IEEE VIS 2024, which won Best Poster Award and remains available as arXiv v1. Project page: https://poloclub.github.io/transformer-explainer/

Journal ref: CHI '26: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 7 (2026), 21 pages

DOI: 10.1145/3772318.3791725

Code: https://github.com/safety-research/circuit-tracer

Project page: https://poloclub.github.io/transformer-explainer

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: Transformer Explainer is "an interactive visualization tool for non-experts to learn Transformers" designed to address the complexity of the Transformer architecture, which "makes it difficult for

Terminology

Summary

Transformer Explainer is an interactive visualization tool for non-experts to learn Transformers designed to address the complexity of the Transformer architecture, which makes it difficult for non-experts to learn. The authors note that existing resources often lack interactivity, rely on static descriptions of simplified architectures, or fail to reflect models’ behavior with real data. To bridge this gap, the tool integrates an overview illustrating the Transformer’s data flow with on-demand explanations that gradually reveal mathematical details, allowing for smooth transitions across abstraction levels to highlight the interplay between high-level structures and low-level operations.

The design of the tool is driven by four primary challenges:

  • C1. Understanding How Input Text Is Processed Across Complex Model Structures: Transformers consist of many repeating blocks, each containing many interacting operations, making it difficult to follow how input data (i.e., token embedding) is transformed into final predictions.

  • C2. Mathematical Complexity in Multi-Head Self-Attention: The mechanism is significantly more complex than other model operations because it involves multiple matrix operations that enable every input text token... to simultaneously interact with every other token.

  • C3. Understanding Hyperparameters’ Impact on Prediction Variability: Many learners are unaware of how the generated probability distribution determines next-token predictions and how hyperparameters shape output variability.

  • C4. Deployment for Scalable Iterative Learning: Hosting a live Transformer model in-browser is a technical challenge due to models being large and computationally intensive.

To address these, the authors established four design goals: G1. Model Overview Prioritizing Token-Centric Data Flow, which uses a Sankey diagram-inspired design to show how input information ‘flows’ through the various components; G2. Visual Disambiguation of Multi-Head Self-Attention with Step-by-Step Visual Explanations using progressive disclosure technique[s]; G3. Dynamic Experimentation Through User-Provided Text and Hyperparameter Manipulation; and G4. Web-based Tool Powered by Live Model for Interactive Learning.

The system features several key components:

  • Overview: The visual design draws inspiration from the Sankey diagram to communicate a high-level overview of how input data flows through the Transformer model. It uses gradient-colored paths and vertical bars scaled to their actual dimensionality to represent token embeddings.

  • Step-by-Step Expanded Explanations: Users can interactively expand the model components through animated transitions to explore mathematical details. This includes the Expanded Self-Attention View, which animates the computation of attention scores in three sequential steps, the Expanded Probabilities View, which incrementally displays each step in probability computation for next-token prediction, and the Expanded Embedding View, which illustrates how each token from the input text is converted into its numerical embedding vector.

  • Real-time Inference and Experimentation: The tool runs a live GPT-2 instance directly in the browser, allowing users to input their own text... and directly manipulate sampling hyperparameters... observing next-token predictions in real time. Users can adjust Temperature, which shapes the generated probability distribution, and Sampling Strategies, such as top-k and top-p.

  • Guided Learning: An interactive, step-by-step text explanation card embedded within the tool that introduces Transformer concepts by following the flow of data.

The effectiveness of Transformer Explainer was evaluated through a 90-participant between-subjects user study comparing the tool against a blog post and an educational video. The findings showed that Transformer Explainer offered significant advantages in improving user understanding and engagement. Specifically, Transformer Explainer participants achieved higher quiz accuracy than Blog (p = 0.021) and Video (p = 0.021) and reported higher achievement of learning objectives than Blog (p = 0.006), and (C) higher learning experience ratings than Blog (p = 0.026) and Video (p = 0.033). Qualitative analysis indicated that interactivity as one of the most helpful features, with users praising the ability to change certain settings and see how temperature, p and k change. Since its launch, Transformer Explainer has attracted over 490,000 users.

Improvements for AI systems

1. Semantic Flow Interpretability Layer

  • What it can do: This improvement integrates a real-time visualization engine directly into the model's inference pipeline. It will map the transformation of token embeddings into high-level semantic trajectories, allowing developers to visually track how specific concepts (e.g., sentiment or subject-verb agreement) are encoded, modified, or lost as data flows through successive Transformer blocks.

2. Multi-Head Semantic Decomposition Interface

  • What it can do: Instead of presenting raw attention matrices, this system will implement a layer that translates multi-head computations into human-readable semantic descriptors. It will automatically categorize and label attention heads by their functional role (e.g., Syntactic Dependency Head, Anaphoric Reference Head, or Contextual Modifier Head), enabling researchers to diagnose exactly which heads are responsible for specific linguistic or logical errors.

3. Probabilistic Sensitivity Dashboard for Hyperparameter Optimization

  • What it can do: This improvement creates a real-time feedback loop between the sampling parameters and the output probability distribution. It will allow developers to observe the Entropy-vs-Creativity tradeoff visually, showing exactly how adjustments to Temperature, Top-k, or Top-p collapse or expand the probability mass across the vocabulary during live inference, facilitating much faster tuning of model personality and coherence.

4. Multi-Tiered Model Transparency API

  • What it can do: This system will provide Abstraction-on-Demand for model auditing. It will allow users (ranging from non-expert auditors to deep-learning researchers) to toggle between different levels of technical depth: high-level architectural summaries for safety compliance, mid-level data flow diagrams for logic verification, and low-level mathematical proofs of specific neuron activations for debugging weight-level biases.

5. Edge-Optimized Diagnostic Probing Engines

  • What it can do: To solve the challenge of deploying large models for interactive learning, this improvement involves developing lightweight, quantized diagnostic twins of large models. These twins will run locally on user hardware to provide real-time, interactive probing of model weights and activations, allowing for rapid, iterative testing of model behavior without the latency or cost of server-side inference.

Sources

Related papers