SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

arXiv:2607.23991 · cs.CL, cs.AI · Submitted 2026-07-27 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-07-27

Updated: 2026-08-30

Comments: EMNLP 26 Main, 27 pages

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements.

Terminology

Abstract

Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM's top-k candidates by combining base logits with both the learned reward signal and a contrastive signal that captures system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.

Sources

Related papers