Can LLMs Reliably Self-Report Adversarial Prefills, and How?
cs.CL
Submitted: 2026-06-22
Updated: 2026-09-01
Comments: EMNLP 2026 (Main)
Code: https://github.com/ngqm/prefill-introspection
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prior work shows that large language models (LLMs) exhibit varying degrees of introspective capability on benign tasks.
Terminology
Abstract
Prior work shows that large language models (LLMs) exhibit varying degrees of introspective capability on benign tasks. We extend the question to safety contexts and examine how reliably a model can recognize that its own prior response was elicited by an adversarial prefill attack. Across ten open-weight instruction-tuned LLMs from 3B to 70B parameters and four safety benchmarks, no model reliably recognizes its own compromised outputs, with models claiming intent on prefilled responses at an average rate of 25.3%. Introspective signal stems primarily from reasoning about safety and refusal. Orthogonalizing models' weights against the refusal direction collapses the gap between claim rates on prefilled and natural outputs to near zero, though the direction is not its unique mediator. Framing the question as internal intention versus external tampering elicits qualitatively different responses on the same models. Training models to mimic correct introspective answers or optimize an introspective objective can improve the accuracy of introspection, but such training does not transfer to the tampering probe and counterintuitively raises attack success rate under adversarial prefill on most models, amounting to a partial mitigation. These findings outline mechanisms underpinning the observed introspective signals in safety contexts and highlight risks in the reliability of LLM self-reports.
Sources
- Eliciting Secret Knowledge from Language Models
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- The Llama 3 Herd of Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Training Agents to Self-Report Misbehavior
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Language Models (Mostly) Know What They Know
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Large Language Models Often Know When They Are Being Evaluated
- Prefill Awareness in Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3 Technical Report
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering