IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies
cs.CL
Submitted: 2026-06-29
Updated: 2026-09-17
Comments: EMNLP 2026 Findings
Code: https://github.com/nxcolelxu/IHDec
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
- Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
- IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
- Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
- Lost in the Middle: How Language Models Use Long Contexts
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency
- Reasoning Up the Instruction Ladder for Controllable Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering