DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-23
Updated: 2026-09-06
Comments: EMNLP 2026 Main Conference
Code: https://github.com/tatsu-lab/alpaca_eval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility.
Terminology
Abstract
Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility. We propose Dynamic SAE Steering for Preference Alignment (DSPA), an inference-time method that makes sparse autoencoder (SAE) steering prompt-conditional. From preference triples, DSPA computes a conditional-difference map linking prompt features to generation-control features; during decoding, it modifies only token-active latents, without base-model weight updates. Across Gemma-2-2B/9B and Qwen3-8B, DSPA improves MT-Bench and is competitive on AlpacaEval while preserving multiple-choice accuracy. Under restricted preference data, DSPA remains robust and can rival the two-stage RAHF-SCIT pipeline while requiring up to 4.47 times fewer alignment-stage FLOPs. Finally, we audit the SAE features DSPA modifies, finding that preference directions are dominated by discourse and stylistic signals, and provide theory clarifying the conditional-difference map estimate and when top- k ablation is principled.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Toy Models of Superposition
- Steering Large Language Model Activations in Sparse Spaces
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- Towards Inference-time Category-wise Safety Steering for Large Language Models
- YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation
- Scaling and evaluating sparse autoencoders
- Language Models are Few-Shot Learners
- BatchTopK Sparse Autoencoders
- PaLM: Scaling Language Modeling with Pathways
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Training language models to follow instructions with human feedback
- PersonalLLM: Tailoring LLMs to Individual Preferences
- Representation Engineering: A Top-Down Approach to AI Transparency
- Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude Perspective
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Gemma 2: Improving Open Language Models at a Practical Size
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks