Overflip: Repetition-Induced Label Flips in Guardrail Models
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 12 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Guardrail models are classifiers deployed to screen malicious prompts and responses in LLM-based services.
Terminology
Abstract
Guardrail models are classifiers deployed to screen malicious prompts and responses in LLM-based services. To meet latency constraints, many lightweight guardrails adopt compact Transformer backbones (e.g., DeBERTa) that are trained with short context windows (typically 512 tokens) and rely on bucketed relative positional encodings to process longer inputs. Prior evaluations assume that a guardrail's decision is stable as the input is lengthened. We show that this assumption can fail. We identify Overflip, a repetition-induced instability where repeating a prompt causes the guardrail's prediction to flip (MAL to BEN) as the sequence grows. We conduct experiments on 9 widely used lightweight guardrail models. Five exhibit MAL to BEN flips on a benchmark of 100 prompts, with confidence margins shrinking steadily with repetition. Among these vulnerable models, flip rates range from 8% to 92%, with first flips occurring at roughly 2.6k--9.4k tokens. Our analysis suggests Overflip differs from traditional attention-dilution baselines, which aim to divert the model's attention away from tokens associated with malicious content, shifting it instead toward unrelated content, such as benign padding or shuffling. While Overflip preserves malicious content, it homogenizes token-level attention over repeated structure and induces a distinct, more gradual attention-dispersion trajectory than padding. Moreover, Overflip poses a greater threat to LLM services than traditional attention dilution methods. Because the bypassed prompt remains semantically intact and is still readily understood by downstream business LLMs, it can transmit malicious intent after passing the guardrail. These findings expose repetition as an attack surface for guardrail models and motivate length-robust evaluation and mitigation.
Sources
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Improve Transformer Models with Better Relative Position Embeddings
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection
- Lost in the Middle: How Language Models Use Long Contexts
- Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection