Many-Tier Instruction Hierarchy in LLM Agents
cs.CL, cs.AI
Submitted: 2026-04-10
Updated: 2026-09-04
Comments: EMNLP 2026 Findings
Code: https://github.com/JHU-CLSP/ManyIHhf.co
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- DeonticBench: A Benchmark for Reasoning over Rules
- IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
- When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
- Kimi K2.5: Visual Agentic Intelligence
- Prompt Injection attack against LLM-integrated Applications
- Generalizing Verifiable Instruction Following
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
- CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
- CCTU: A Benchmark for Tool Use under Complex Constraints
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
- Jailbreak Distillation: Renewable Safety Benchmarking
- Effective Prompt Extraction from Language Models
- IHEval: Evaluating Language Models on Following the Instruction Hierarchy
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering