On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study
cs.CL
Submitted: 2026-06-10
Updated: 2026-09-10
Comments: Published as a workshop paper at BlackBoxNLP 2026
Code: https://github.com/apple/ml-act
License: http://creativecommons.org/licenses/by/4.0/
The gist: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive.
Terminology
Abstract
Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet previously overlooked interaction with the training paradigm: activation steering methods are far less effective on instruction-tuned models than on their base counterparts. Simple prompting and full-fledged supervised fine-tuning, on the other hand, are viable options for concept injection, but are not as good at concept removal. Finally, cheaply computed textual metrics highly correlate to costly LLM-as-judge scores, and provide insights on the behavior of conditioning methods.
Sources
- ExpertLens: Activation steering features are highly interpretable
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Curious Case of Neural Text Degeneration
- Mistral 7B
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Steering Llama 2 via Contrastive Activation Addition
- LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
- Perplexity Cannot Always Tell Right from Wrong
- Neural Text Generation with Unlikelihood Training
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- ReFT: Representation Finetuning for Language Models
- Qwen3 Technical Report
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering