iPOE: Interpretable Prompt Optimization via Explanations
cs.CL
Submitted: 2026-05-18
Updated: 2026-09-07
Comments: published at EMNLP 2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Prompt optimization has often been framed as a discrete search problem to find high-performing and robust instructions for a large language model (LLM).
Terminology
Abstract
Prompt optimization has often been framed as a discrete search problem to find high-performing and robust instructions for a large language model (LLM). However, the search result might not make it transparent why and where specific prompt changes lead to performance gains. This is in contrast to how humans are instructed for annotation tasks. Here, researchers carefully design annotation guidelines, leading to enhanced annotation consistency. Our paper aims at joining these two approaches and introduces iPOE, a novel interpretable prompt optimization strategy via explanations. We guide the prompt optimization process by automatically created guidelines from explanations of annotation decisions (either automatically generated or from humans). This set of guidelines is furthermore optimized by a series of operations, including removing, adding, shuffling, and merging. The resulting prompt includes guidelines that instruct the annotation, making the decision process of the LLM and the optimization transparent. It therefore supports also laypeople in prompt optimization. In our experiments on four datasets, we find that iPOE can improve over the evaluated baselines by up to 39% and LLM explanations can replace human explanations in the proposed method. Moreover, our interpretability validation study demonstrates that humans and LLMs substantially agree on which guidelines contribute to their annotations, achieving a Cohen's kappa score between annotators and LLMs of up to 0.68, 0.80, and 0.95 for the emotion classification, medical fact-checking, and hate speech detection tasks respectively.
Sources
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
- The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink
- Qwen3 Technical Report
- InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction
- Large Language Models as Optimizers
- PromptNER: A Prompting Method for Few-shot Named Entity Recognition via k Nearest Neighbor Search
- Large Language Models Are Human-Level Prompt Engineers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering