STR: Supervised Transcoder Replacement for Reducing Steering Side Effects
cs.AI, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Transcoders Find Interpretable LLM Feature Circuits
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Forecasting Side Effects of Activation Steering
- Discovering Language Model Behaviors with Model-Written Evaluations
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- A StrongREJECT for Empty Jailbreaks
- Steering Language Models With Activation Engineering
- Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
- Qwen3 Technical Report
- Representation Engineering: A Top-Down Approach to AI Transparency
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection