On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study

arXiv:2606.12234 · cs.CL · Submitted 2026-06-10 · Read on arXiv

cs.CL

Submitted: 2026-06-10

Updated: 2026-09-10

Comments: Published as a workshop paper at BlackBoxNLP 2026

Code: https://github.com/apple/ml-act

License: http://creativecommons.org/licenses/by/4.0/

The gist: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive.

Terminology

Abstract

Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet previously overlooked interaction with the training paradigm: activation steering methods are far less effective on instruction-tuned models than on their base counterparts. Simple prompting and full-fledged supervised fine-tuning, on the other hand, are viable options for concept injection, but are not as good at concept removal. Finally, cheaply computed textual metrics highly correlate to costly LLM-as-judge scores, and provide insights on the behavior of conditioning methods.

Sources

Related papers