Evaluating Steering Techniques using Human Similarity Judgments
cs.AI
Submitted: 2025-05-25
Updated: 2026-09-04
License: http://creativecommons.org/licenses/by/4.0/
The gist: Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition.
Terminology
Abstract
Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. Using a well-established triadic similarity judgment task, we assessed steered LLMs on their ability to flexibly judge similarity between concepts based on size or kind, two central dimensions organizing human mental representations. We found that prompt-based steering methods outperformed other methods both in terms of steering accuracy and model-to-human alignment. We also found LLMs were biased towards `kind' similarity and struggled with `size' alignment. This evaluation approach, grounded in human cognition, adds further support to the efficacy of prompt-based steering and reveals privileged representational axes in LLMs prior to steering.
Sources
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Scaling and evaluating sparse autoencoders
- In-Context Learning Creates Task Vectors
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex
- Aligning Large Language Models with Human Preferences through Representation Engineering
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Function Vectors in Large Language Models
- Getting aligned on representational alignment
- Adaptively Learning the Crowd Kernel
- Steering Language Models With Activation Engineering
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection