Steering Language Model Goals with Value Transplant
cs.LG, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/pat-jj/value-transplant
Terminology
Sources
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Alignment faking in large language models
- How's it going? Reinforcement learning in language models recruits a functional welfare axis
- Risks from Learned Optimization in Advanced Machine Learning Systems
- The Value Axis: Language Models Encode Whether They're on the Right Track
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- Understanding and Controlling a Maze-Solving Policy Network
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Steering Language Models With Activation Engineering
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks