Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
Jiwon Moon, Yerin Hwang, Kyomin Jung
cs.CL
Submitted: 2026-08-21
Updated: 2026-08-24
Comments: Accepted to EMNLP 2026 (Main). Code and data are available at https://github.com/g1moon/Language-Shapes-IH
Code: https://github.com/g1moon/Language-Shapes-IH
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
- Qwen Technical Report
- ALLaM: Large Language Models for Arabic and English
- K-EXAONE Technical Report
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- The Llama 3 Herd of Models
- IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
- Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
- Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
- XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
- Ministral 3
- SysBench: Can Large Language Models Follow System Messages?
- EuroLLM-22B: Technical Report
- OpenAI GPT-5 System Card
- HyperCLOVA X THINK Technical Report
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
- Qwen3 Technical Report
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering