The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
cs.CL, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/Usama1002/system-prompt-illusion-cka
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Refusal in Language Models Is Mediated by a Single Direction
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- InternLM2 Technical Report
- Reliability of CKA as a Similarity Measure in Deep Learning
- Transformer Feed-Forward Layers Are Key-Value Memories
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
- In-Context Learning Creates Task Vectors
- Mistral 7B
- Scaling Laws for Neural Language Models
- Similarity of Neural Network Representations Revisited
- Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Gemma: Open Models Based on Gemini Research and Technology
- Correcting Biased Centered Kernel Alignment Measures in Biological and Artificial Neural Networks
- In-context Learning and Induction Heads
- Training language models to follow instructions with human feedback
- Nemotron-4 15B Technical Report
- Where does an LLM begin computing an instruction?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering