HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix
cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/meta-llama/llama-models
Terminology
Sources
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
- Gemma 3 Technical Report
- LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
- D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
- Mechanistic Analysis of Alignment Algorithms in Language Models
- LLM Agents Already Know When to Call Tools -- Even Without Reasoning
- Steering Language Models With Activation Engineering
- Qwen3 Technical Report
- Rep2Text: Decoding Full Text from a Single LLM Token Representation
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering