Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026
Code: https://github.com/DecayingSeart/CREST
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) finetuned for specialized domains represent crucial high-impact applications.
Terminology
Abstract
Large language models (LLMs) finetuned for specialized domains represent crucial high-impact applications. Inference-time alignment improves safety degraded from specialization finetuning without requiring substantial computational resources, complementing finetuning-based methods with an easy-to-use, plug-and-play solution. However, existing inference-time methods fail to reliably improve safety without disrupting domain capability. We identify the root cause as complementary expertise orthogonality: specialized base models and general-domain guidance models have orthogonal competencies, making the guidance signal unreliable for specialized generation. This primarily manifests as stop token interference, where the guidance model's tendency toward continuation overrides the base model's decision to stop, burying correct answers under guidance-induced continuation. To address this problem, we propose CREST, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely. CREST improves safety where specialization has weakened it while preserving both domain-specific capability and the safety of already well-aligned models, outperforming baselines by up to 22.2% on safety benchmarks. Our code is available at: https://github.com/DecayingSeart/CREST.
Sources
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- The Llama 3 Herd of Models
- Qwen2.5-Coder Technical Report
- MedGemma Technical Report
- Locking Down the Finetuned LLMs Safety
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering