SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding
cs.CL, cs.AI
Submitted: 2026-07-27
Updated: 2026-08-30
Comments: EMNLP 26 Main, 27 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements.
Terminology
Abstract
Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM's top-k candidates by combining base logits with both the learned reward signal and a contrastive signal that captures system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.
Sources
- The Llama 3 Herd of Models
- A Closer Look at System Prompt Robustness
- DeepSeek-V3 Technical Report
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Instruction-Following Evaluation for Large Language Models
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering