Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/abovefiramament/DHSA
Terminology
Sources
- The Llama 3 Herd of Models
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Decoupled Weight Decay Regularization
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Qwen2.5 Technical Report
- Activation Steering via Generative Causal Mediation
- Steering Language Models With Activation Engineering
- Attention Is All You Need
- Operationalising the Superficial Alignment Hypothesis via Task Complexity
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Qwen2 Technical Report
- LIONs: An Empirically Optimized Approach to Align Language Models
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks