STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: Accepted to Findings of the Association for Computational Linguistics: AACL-IJCNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Textual-gradient methods automate prompt optimization through natural-language feedback, but their iterative updates can be unstable.
Terminology
Abstract
Textual-gradient methods automate prompt optimization through natural-language feedback, but their iterative updates can be unstable. We identify two sources of this instability: noisy gradients produced from already-correct examples and over-specialization to hard cases that degrades performance on simpler inputs. We introduce STEVE, a stabilization framework with two coupled mechanisms. Error-Driven Refinement generates gradients only from incorrectly handled examples, concentrating updates on informative failures. Regularized Verification treats every update as provisional and accepts it only when improvement on hard cases does not cause unacceptable regression on a preservation set. Across ten reasoning benchmarks, three evaluator/optimizer models, and established prompt-optimization baselines, STEVE reduces degradation and produces more robust prompts. Additional evaluations with gpt-5.4-mini/gpt-5.4 on symbolic reasoning, GSM8K-Platinum, and DS-1000 show that these gains persist with newer models and larger test sets. STEVE therefore provides a practical way to improve the stability and effectiveness of textual-gradient prompt optimization.
Sources
- Training Verifiers to Solve Math Word Problems
- Introducing MAPO: Momentum-Aided Gradient Descent Prompt Optimization
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning
- Language Model Cascades
- Measuring Massive Multitask Language Understanding
- Solving General Arithmetic Word Problems
- GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models
- Prioritized Experience Replay
- The Power of Scale for Parameter-Efficient Prompt Tuning
- Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
- Survival of the Safest: Towards Secure Prompt Optimization through Interleaved Multi-Objective Evolution
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
- Do Large Language Model Benchmarks Test Reliability?
- Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
- SPRIG: Improving Large Language Model Performance by System Prompt Optimization
- PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization
- Survival of the Most Influential Prompts: Efficient Black-Box Prompt Search via Clustering and Pruning
- GPS: Genetic Prompt Search for Efficient Few-shot Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering