Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

arXiv:2608.09900 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Martínez, Pedro Reviriego

University of the Ryukyus · Technion · Universidad Politécnica de Madrid

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: Decoding-Level Taboo is a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime to force large language models (LLMs) off their nominal decoding paths.

Terminology

Summary

Decoding-Level Taboo is a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime to force large language models (LLMs) off their nominal decoding paths. The method dynamically masks the top-i candidate tokens at word-initial boundaries, forcing the model into machine circumlocution by constructing semantically valid, off-path alternatives. The severity of the stress is quantified via injected surprisal (∆S), the extra bits of surprisal in the model's own belief when emitting the forced alternative instead of its preferred token.

The paper evaluates Taboo across four open-weight model families (Qwen2.5, Gemma-3, Llama-3, OLMo-2) and four benchmarks (GSM8K, MMLU, TriviaQA, HumanEval). Key findings include:

  1. Off-path robustness is governed jointly by scale and post-training alignment: Aligned checkpoints absorb a larger effective perturbation yet retain far more of their multi-step reasoning. This effect widens with scale, is family-dependent (absent in Llama-3 at 7–8B, though it reappears by 70B), task-specific (generative reasoning, not MCQ or factual recall), and bounded by formal syntax (HumanEval collapses to zero).

  2. Parameter scale generally improves off-path robustness on generative reasoning, though not monotonically: On Qwen models for GSM8K, conditional retention at the mildest dose (i=1) rises with size for base checkpoints overall (6% at 0.5B to 59% at 32B), though 7B base dips to 8%. Aligned counterparts are markedly more robust, from 6% at 0.5B to 93% at 32B and 85% at 72B. The largest base model (72B) retains less than 32B-base (39% vs. 59% at i=1), showing raw scale alone does not guarantee off-path stability.

  3. Instruction tuning confers off-path robustness, but only for multi-step reasoning and only in some families: On GSM8K, aligned checkpoints retain a far larger fraction of baseline-correct items under Taboo. At i=1, conditional retention rises from 8% to 48% in Qwen2.5-7B, 15% to 45% in OLMo-2-7B, and 25% to 90% in Gemma-3-12B. Llama-3.1-8B shows essentially no separation between base and instruct (14% vs. 16% at i=1). The benefit is task-specific: on TriviaQA the base–instruct gap is small, while on MMLU both checkpoints are weak and non-monotone (an MCQ-floor artifact).

  4. Domain constraints matter: TriviaQA exhibits the highest global retention due to natural language surface-form flexibility; MMLU collapses toward the guessing floor because the single-letter answer is inherently word-initial; GSM8K demonstrates moderate retention with eventual logic drift; HumanEval serves as a negative control where masking reserved keywords at word boundaries instantly breaks Python AST construction.

  5. Mechanistic insight via injected surprisal: In three families (Qwen, OLMo, Gemma), aligned checkpoints absorb more injected surprisal per intervention than their base counterparts (sharper next-token distributions), yet retain more. Llama-3 is the exception, absorbing no more surprisal than its base at both 8B and 70B, yet at 70B alignment recovers a large retention gain (36% → 75% at i=1), suggesting a mechanistically distinct route to robustness not passing through distribution sharpening.

The paper also discusses broader applications: auditing alignment and surface-level refusals, generating diverse trajectories for synthetic CoT datasets, zero-training stress testing for structured generation, and taboo-guided alignment for improving off-path robustness. Limitations include the scope of intervention regimes, English-only evaluation, restricted model architectures, and mixed precision constraints at larger scales.

Improvements for AI systems

Improvements to AI systems:

  1. Add a runtime “off-path stress test” module that dynamically masks top-1 to top-i tokens at word boundaries during inference. This forces the model to generate alternative valid continuations, revealing hidden brittleness in multi-step reasoning before deployment. The improved system can self-audit its own robustness under adversarial decoding conditions without retraining.

  2. Implement “injected surprisal” (∆S) as a calibration metric during alignment fine-tuning. By measuring how much surprisal an aligned model absorbs per forced alternative, the system can detect whether its robustness comes from sharper next-token distributions or from a distinct mechanism (as in Llama-3-70B). The improved system can automatically adjust its decoding strategy (e.g., soften or sharpen logits) based on task type and model family to maximize retention under perturbation.

  3. Introduce “taboo-guided alignment” as a post-training step: after instruction tuning, run Taboo interventions on GSM8K-style generative reasoning tasks, and use the retention gap (base vs. aligned) as a reward signal to further fine-tune the model. The improved system can explicitly learn to maintain logical coherence even when forced to avoid its preferred token sequences, improving off-path robustness for multi-step math and code reasoning.

  4. Add a “domain-aware masking policy” that adapts the intervention severity based on surface-form flexibility. For generative reasoning (GSM8K), use moderate masking (i=1–3) to test logic drift; for MCQ (MMLU), avoid word-initial masking of single-letter answers to prevent floor artifacts; for code (HumanEval), skip masking reserved keywords to preserve syntax. The improved system can automatically select the appropriate stress regime per task, enabling zero-training stress testing for structured generation without breaking formal constraints.

  5. Build a “scale-aware robustness predictor” that uses the observed family-dependent, non-monotonic scaling curves (e.g., Qwen 7B dip, 72B base regression) to predict which checkpoint size and alignment stage will yield the best off-path retention for a given task. The improved system can recommend the optimal model variant (e.g., 32B aligned over 72B base for GSM8K) to avoid wasted compute and ensure stable performance under unexpected input perturbations.

  6. Enable “diverse trajectory generation” for synthetic CoT datasets by using Taboo to force multiple off-path reasoning paths for the same problem. The improved system can generate richer, more varied chain-of-thought training data that includes non-nominal but valid solutions, enhancing the model’s ability to generalize to novel reasoning strategies and reducing overfitting to preferred token sequences.

  7. Add a “refusal and alignment audit” feature that applies Taboo to safety-critical prompts to test whether alignment is superficial (surface-level refusals) or deeply embedded. The improved system can detect if a model’s refusal behavior collapses under forced alternative phrasing, and then trigger additional safety fine-tuning or a fallback to a more robust aligned checkpoint.

Sources

Related papers