Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs
cs.CR, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/lukasbruna/output-prefix-attack
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Training Large Language Models to Reason in a Continuous Latent Space
- OverThink: Slowdown Attacks on Reasoning LLMs
- Reinforcement Learning from Human Feedback
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- Prefill Awareness in Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs