CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad, Ali Abbasi, Seyedarmin Azizi
cs.LG
Submitted: 2026-08-06
Code: https://github.com/mehrshad-sdtn/CircuitSteer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
Terminology
Abstract
Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co-activation and the geometric alignment of decoder directions, we isolate the specific multi-layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi-point interventions to guide the model's internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion-intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency-preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single-point interventions. Code is available at https://github.com/mehrshad-sdtn/CircuitSteer.
Sources
- TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- Mechanistic Permutability: Match Features Across Layers
- Training Verifiers to Solve Math Word Problems
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models
- Interpretable LLM Guardrails via Sparse Representation Steering
- Editing Models with Task Arithmetic
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- Steering Llama 2 via Contrastive Activation Addition
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Towards Understanding Sycophancy in Language Models
- Analyzing the Generalization and Reliability of Steering Vectors
- NeuroMambaLLM: Dynamic Graph Learning of fMRI Functional Connectivity in Autistic Brains Using Mamba and Language Model Reasoning
- Steering Language Models With Activation Engineering
- Improving LLM Reasoning through Interpretable Role-Playing Steering
- Exploring Representations and Interventions in Time Series Foundation Models
- Small Graph Is All You Need: DeepStateGNN for Scalable Traffic Forecasting
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks