From Monolithic to Modular: Segment-level Automatic Prompt Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From Monolithic to Modular: Segment-level Automatic Prompt Optimization".
Jane: The paper was written by Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov et al. from ITMO University and AXXX.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back to the show. Today we are cracking open a fresh one from the arXiv pile, and it's called "From Monolithic to Modular: Segment-level Automatic Prompt Optimization." Jane, I gotta say, that title alone tells a story.
Jane: It really does, Tom. And it's a story about how we talk to these big language models. For a long time, when we wanted a model to do something, we'd write one big block of instructions, a monolithic prompt, and hope it worked.
Tom: Right, and if it didn't work, you'd just rewrite the whole thing and hope again. That's the "monolithic" part.
Jane: Exactly. This paper from the folks at ITMO University and AXXX is saying, hey, let's stop treating the prompt like a single block of text. Let's break it down into parts.
Tom: And those parts, they've got a nice clean taxonomy. You've got the role, like "you are a helpful assistant." Then the context, which is the background info. The tasks, which are the actual instructions. And finally, the output format.
Lu: It's a really natural decomposition when you think about it. As a researcher, I see this as moving from treating the prompt as an opaque string to treating it as a structured object with distinct, addressable components.
Meng: And from my side, that's a huge deal for practical work. If I know the output format section is causing problems, I want to fix just that, not risk breaking the whole thing by rewriting everything.
Jane: That's the core idea. They call it SAPO, Segment-level APO. Instead of one big rewrite, the system figures out which segment is weak and which segments are strong, and it only touches the weak ones.
Tom: So it's like fixing a car engine. You don't replace the whole engine just because the spark plugs are bad. You diagnose the problem, find the faulty part, and swap it out.
Jane: That's the analogy, Tom. And the promise is that you get better results without the side effects of a full rewrite, which has been a real problem in this field.
Tom: A problem they call "prompt drifting," right? You fix one thing and break two others.
Lu: Precisely. The paper is very explicit about that failure mode. It's a persistent limitation of the older, monolithic approaches.
Meng: So the title is really a thesis statement. It's saying we need to move away from the old way of doing things. I'm curious to see how they actually pull it off in the method.
Jane: That's the next step. We've got the big idea, now let's see the machinery under the hood. That's coming up next.
Summary: Tom: So we've established that "From Monolithic to Modular: Segment-level Automatic Prompt Optimization" is about being surgical with our edits to prompts. Jane, what's the actual summary of how they do it?
Jane: Okay, so picture this. You start with a basic prompt. The system, SAPO, runs that prompt on a training set of examples. It then looks at the five best-performing examples and the five worst-performing examples.
Tom: The top five and the bottom five. That's the contrastive evidence.
Jane: Exactly. It shows those examples to the language model and asks it to figure out which parts of the prompt, which segments, are responsible for the good results and which ones are causing the bad results.
Lu: It's a form of diagnosis. The model is acting as a detective, attributing success and failure to specific parts of the instructions.
Meng: And this is all done with the same model you're optimizing for, right? It's a self-contained loop.
Jane: Yes, Meng. It uses one LLM for everything. It segments the prompt, it analyzes the evidence, and it generates new candidates. The key constraint is that when it generates a new prompt, it's told to keep the strong segments exactly as they are and only revise the weak ones.
Tom: So it's not just generating random variations. It's generating targeted fixes.
Jane: Right. And then it tests those new candidates on a separate validation set. If a candidate scores better than the current prompt, it gets accepted. If not, it keeps the old one. It's a very conservative, gated process.
Meng: That validation gate is crucial. It stops the optimizer from going off the rails and making things worse. It's a safety mechanism.
Lu: And there's a clever tie-break rule too. If two candidates score the same, it picks the one that's most similar to the current prompt. That's a strong prior for minimal change.
Tom: So it's not just about getting a higher score. It's about getting a higher score with the smallest possible intervention.
Jane: That's the summary of the loop. And they tested it on a bunch of different tasks. We're talking question answering, summarization, math problems, sentiment analysis, commonsense reasoning.
Meng: That's a solid spread. It's not just one type of task, which makes the results more believable.
Tom: And the results, Jane? What did they find?
Jane: They found that SAPO beat the other automatic prompt optimization methods on average, on both GPT-three point five-Turbo and GPT-4o-mini. The gains were especially big on the math dataset, GSM8K.
Lu: Which makes sense. Math problems are very sensitive to the exact format of the instructions and the expected output. Segment-level control really shines there.
Meng: So the summary is: a self-diagnosing, self-correcting loop that makes small, targeted changes and only accepts them if they're proven to work. That's a really clean design.
Jane: It is. And now we need to talk about what this actually means for how we build with these models. That's the next part of our discussion.
Improvements: Tom: We've covered the "what" of "From Monolithic to Modular: Segment-level Automatic Prompt Optimization." Now let's talk about the "so what." What are the real improvements here?
Jane: I think the biggest improvement is the shift in mindset. We're moving from hoping a model understands us to systematically engineering how it understands us.
Lu: And that's a profound shift. The paper shows that you can get a +five point one three percent average gain over the best baseline on GPT-three point five-Turbo and a +seven point two five percent gain on GPT-4o-mini. Those aren't trivial numbers.
Meng: The gains are one thing, but for me, the improvement is in the reliability. The fact that the optimization is monotone, meaning it only ever accepts a change if it improves validation performance, that's a massive practical improvement.
Tom: So you're saying it's not just about being better, it's about being safer to use.
Meng: Exactly. In a production system, you can't have a prompt optimizer that randomly makes things worse. This gives you a guarantee that you won't regress.
Jane: And that's tied to the segment-level control. By preserving the strong segments, you're actively preventing the "drifting" problem we talked about earlier.
Lu: It also makes the optimization process interpretable. After it's done, you can look at the final prompt and see exactly which segments were changed and why. You can't do that with a monolithic rewrite.
Tom: That's a good point. It's not a black box anymore. You can audit the changes.
Meng: And that auditability is what makes me think this could be adopted in industry. It's not just a cool research trick; it's a tool that can be integrated into a workflow.
Jane: And the paper even shows the ablation study, where they remove or revert each segment. It confirms that the task segment is the most important, and the context segment is the least sensitive. That's actionable knowledge.
Tom: So the improvement isn't just a better score. It's a more robust, more understandable, and more practical way to do prompt engineering.
Lu: I'd say it's a step towards making prompt optimization a proper engineering discipline rather than an art form.
Meng: And that's something I can get behind. It gives me a tool that I can trust to not break things.
Jane: So we've got a method that's more precise, more reliable, and more transparent. That's a strong package. But what does this mean for the bigger picture? That's what we should tackle in our final thoughts.
Conclusion: Tom: Alright, we're wrapping up our time with "From Monolithic to Modular: Segment-level Automatic Prompt Optimization." Jane, give us the final take.
Jane: The take is that this paper gives us a smarter way to talk to AI. Instead of throwing the whole prompt away and starting over, it teaches the system to fix the broken part while keeping what works.
Lu: And it does so with a clear, measurable methodology. The results across five different datasets and two different models show that this isn't a fluke. It's a reliable improvement.
Meng: For me, the biggest win is the safety. The validation gate and the conservative tie-break mean I can run this optimizer and not worry about it breaking my application. That's what makes it production-ready.
Tom: It's a shift from guesswork to engineering. And that's a shift we should all be excited about.
Jane: We're saying goodbye to this paper, but we're taking its ideas with us. The idea that structure and modularity are the keys to better AI interaction is a powerful one.
Lu: It opens the door to more complex prompt structures and even adaptive systems that can discover new segments on their own.
Meng: And that future work is exactly what I'd want to see next. More automation, more adaptability.
Tom: Well said, everyone. That's all for this paper. Thanks for listening, and we'll be back soon with another exciting piece of research to break down. See you then.
Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov, Daniil Sukhorukov, Ekaterina Averkova
ITMO University · AXXX
cs.AI, cs.CL, cs.LG
Submitted: 2026-07-21
Updated: 2026-08-13
Comments: Accepted at the IJCAI-ECAI 2026 Workshop on Robustifying Generative AI for Reliable, Safe, and Human-Centric Systems (RobustifAI)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 57/100
The gist: The paper presents SAPO, a segment-level Automatic Prompt Optimization (APO) method that addresses a persistent limitation in existing APO pipelines: they "often rewrite prompts monolithically, which
Key concepts
- Monolithic Prompt
- A traditional approach where a single, large block of instructions is used to guide an AI model. If this monolithic prompt fails, the entire block must be manually rewritten, which often leads to unintended side effects or 'prompt drifting' where fixing one part breaks another part.
- Segment-level Automatic Prompt Optimization (SAPO)
- A method that breaks down a prompt into distinct parts like role, context, tasks, and output format. SAPO then uses the LLM itself to identify which specific segments are causing poor performance. It only revises these weak segments while keeping the strong ones untouched.
- Prompt Drifting
- A persistent problem where an attempt to fix a single issue within a monolithic prompt causes unintended negative consequences elsewhere in the original instructions. SAPO aims to prevent this by making surgical, targeted edits rather than whole-prompt rewrites.
Terminology
Summary
The paper presents SAPO, a segment-level Automatic Prompt Optimization (APO) method that addresses a persistent limitation in existing APO pipelines: they often rewrite prompts monolithically, which can improve one behavior while degrading others.
SAPO decomposes prompts into four explicit segments—role, context, tasks, and output format—then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation.
The authors describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals.
Using evaluation across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.
The paper motivates the work by noting that Large language models (LLMs) are increasingly used as general-purpose interfaces for NLP tasks, including instruction following, reasoning, and generation.
In practical settings, adaptation has shifted from finetuning toward prompt design, where system behavior is controlled through natural-language instructions.
Consequently, prompt quality becomes a primary reliability bottleneck: small wording changes can produce large behavioral shifts, while manual prompt iteration remains costly and unstable.
The authors identify a persistent limitation: "many APO pipelines still optimize prompts as monolithic strings. In practice, prompts are compositional artifacts: role framing, context grounding, task directives, and output-format constraints contribute differently to downstream behavior. Prior work reports
prompt drifting and instability in iterative optimization, where edits that fix one subset of cases may degrade previously correct behavior. The core challenge is
not merely generating more prompt variants, but controlling where and how edits are applied."
SAPO contributes three key elements:
-
Segment-explicit optimization objective:
prompts are decomposed into role, context, tasks, and output format, enabling targeted rather than monolithic edits.
-
Constrained candidate synthesis and selection:
candidates preserve strong segments, revise weak segments, and use an edit-distance tie-break for conservative updates.
-
Contrastive diagnostic stage:
update decisions are grounded in top/bottom evidence, including class-aware handling for discrete per-example metrics.
The evaluation covers five datasets with different task types: SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K, on two model backbones: GPT-3.5-Turbo and GPT-4o-mini. SAPO achieves the best average score against APE, OPRO, EvoPrompt, GEPA, and StraGO, with relative gains of +5.13% on GPT-3.5-Turbo and +7.25% on GPT-4o-mini
compared to the strongest competing baseline by average score.
The paper reviews search-based APO methods: APE optimizes instruction candidates via generation and selection; ProTeGi applies textual gradients with beam search and bandit-style selection; OPRO and EvoPrompt extend search with trajectory-conditioned optimization and evolutionary operators; GEPA introduces reflective prompt evolution. The authors note that these methods primarily operate as global prompt-level rewrites rather than as explicit segment-constrained updates.
The paper also reviews structured and modular optimization trends: DSPy and TextGrad cast optimization in programmatic or graph-based forms; Promptomatix emphasizes modular orchestration and cost-aware refinement; section-local optimization explicitly operates over fixed prompt components and reports improved robustness in small-model settings.
The gap summary states: "Existing literature provides either strong search performance or improved robustness diagnostics... but there is still limited evidence on simple black-box pipelines that jointly enforce segment-level controllability, preserve known-strong prompt components, and remain implementation-light under a unified multi-task protocol."
SAPO's positioning: "SAPO sits at the intersection of static meta-prompting and structure-aware optimization. It preserves the deployment simplicity of black-box APO, while introducing explicit segment-level diagnosis and constrained updates that preserve strong components and localize revisions."
The paper formalizes the optimization objective. Given an initial prompt template P(0) with placeholder input, datasets are split into training and validation parts. Given one LLM M and task-dependent metric Qτ, the objective is: P* = arg max Qτ(P; Dval, M). At iteration t, the optimizer applies an update operator U to the current prompt: P(t+1) = U(P(t), Dtrain, Dval). The operator is acceptance-constrained: "if the best candidate generated at iteration t does not improve validation quality, the prompt is kept unchanged. This converts the procedure into a monotone, validation-gated search over prompt space, which directly targets robustness against destructive rewrites."
Prompt structure is represented as four explicit segments: S(P) = srole, scontext, stasks, soutput format.
The optimization loop consists of five steps:
-
Decompose the current prompt into segments role, context, tasks, output format.
-
Run Stage A: evaluate on train, extract top-5/bottom-5 evidence, infer weak/strong segments and recommendations.
-
Run Stage B: generate K constrained prompt candidates.
-
Evaluate candidates on validation; choose the best by score, with edit-distance tie-break.
-
Accept the candidate only if validation score improves; otherwise keep the current prompt.
The four segments map to distinct and operationally separable control dimensions in instruction-based LLM use
:
-
Role (srole):
defines behavioral stance (e.g., classifier, summarizer, analyst)
and canmaterially affect downstream behavior.
-
Context (scontext):
encodes task grounding, input injection, and domain constraints (including input placement). It controls what information is available and how the model conditions on it.
-
Tasks (stasks):
specifies actionable requirements and decision rules. It is the main locus for correcting underspecified or ambiguous instructions.
-
Output format (soutput format):
governs the response schema and formatting constraints. It is critical in tasks where metric outcomes depend on strict label/output conventions.
These four segments provide a compact decomposition that is expressive enough for heterogeneous NLP tasks while remaining small enough for stable, low-cost iterative optimization.
Stage A: Evidence extraction and segment-level diagnosis. At iteration t, the current prompt P(t) is evaluated on Dtrain: ŷi = M(P(t), xi), qi = Qτ(P(t); (xi, yi), M). Examples are ranked by qi, and two contrastive evidence sets are extracted: Bgood = Top-5(qi), Bbad = Bottom-5(qi). For discrete per-example metrics (e.g., ExactMatch with qi ∈ 0, 1), the method uses "class-aware evidence selection: Bgood is drawn first from positive examples (qi = 1), and Bbad is drawn first from negative examples (qi = 0). If one side has fewer than five examples, the remainder is backfilled from the global rank order. Given the evidence sets and segment decomposition,
one LLM with static meta-prompts infers structured diagnostic outputs: weak segments, strong segments, and recommendations. Intuitively, strong segments are associated with consistently successful evidence, while weak segments are associated with failure cases and become primary targets for revision."
Stage B: Candidate synthesis. The same LLM receives current prompt, segment decomposition, weak/strong labels, and recommendations, then generates K improved prompt candidates. Candidate synthesis is explicitly constrained to preserve segments listed in strong segments and primarily modify segments listed in weak segments, reducing cross-segment interference.
Each candidate is evaluated on Dval. The best candidate is accepted only if it improves the current validation score. Under score ties, the method selects the candidate with the smallest edit distance to the current prompt,
which favors minimal edits and helps preserve validated prompt behavior.
Datasets and task coverage: SQuADv2 (extractive QA), TweetEval (social NLP classification), XSUM (abstractive summarization), CommonGen (constrained commonsense generation), and GSM8K (mathematical reasoning). Metrics: BERTScore F1 for XSUM, CG, SQuAD2; F1 for TweetEval; ExactMatch for GSM8K.
Models: GPT-3.5-Turbo and GPT-4o-mini.
Data split and protocol: "Each dataset is split into train/validation/test with sizes 150/100/all. The train split is used for top/bottom evidence extraction, validation is used for candidate selection during optimization, and the held-out test split is used for final reporting. This separation is strictly enforced for all iterative methods to reduce selection leakage."
Generation settings and baselines: K = 5 candidates per iteration, sampling temperature of 0, fixed across all experiments and identical for SAPO and all iterative APO baselines. Baselines include Zero-shot prompting, APE, OPRO, EvoPrompt, GEPA, and StraGO.
Runs and reporting: Each method/model setup is run five times with different random seeds and data samples. Main tables report mean values across runs.
On GPT-3.5-Turbo: SAPO achieves the highest average score. It yields gains on SQuAD2, GSM8K, CommonGen, and TweetEval, while trailing GEPA on XSUM. The largest relative gain appears on GSM8K. Specific scores: SQuAD2 0.922, GSM8K 0.835, CG 0.911, TweetEval 0.575, XSum 0.852, Average 0.819.
On GPT-4o-mini: SAPO again delivers the highest average score and outperforms all compared baselines on all five datasets.
Specific scores: SQuAD2 0.931, GSM8K 0.952, CG 0.875, TweetEval 0.593, XSum 0.789, Average 0.828.
The average gain relative to the best competing baseline by average score is +5.13% on GPT-3.5-Turbo and +7.25% on GPT-4o-mini.
Why segment-level updates help: "The observed gains are consistent with the hypothesis that explicit weak/strong segment separation reduces destructive interference between edits. Instead of globally rewriting prompts, SAPO localizes revisions to weak components while preserving high-performing structure."
When improvements are likely: Improvements are most pronounced when tasks depend on strict output conventions or multi-part instructions. In such settings, preserving strong format/context segments appears as important as improving task wording.
Limitations: A key limitation of this work is that SAPO uses a fixed and relatively small segment set. While this decomposition is practical and interpretable, it may not capture finer-grained prompt structure needed for some tasks.
Future work should explore richer segment taxonomies and adaptive segment discovery, including LLM-based automatic generation of new segment types during optimization.
Additionally, our current setup uses fixed top-5/bottom-5 evidence windows, which may miss finer-grained error modes.
The ablation study quantifies the contribution of individual prompt segments. All ablation families are negative on average, meaning the final SAPO prompt remains the strongest configuration overall.
Key findings:
-
Reverting tasks to initial wording causes the largest degradation (Avg Δ = −0.1038), with the strongest drop on GSM8K (−0.2800), confirming
task-level refinement is the main driver of improvements.
-
Removing response-format constraints hurts on average (Avg Δ = −0.0354).
-
Role removal is unfavorable overall (Avg Δ = −0.0291).
-
Context reversion is the least harmful perturbation (Avg Δ = −0.0031), suggesting
context edits are comparatively stable, while tasks and output constraints are the most sensitive levers.
The paper concludes: "We presented SAPO, a modular APO framework that replaces monolithic rewrites with segment-level optimization. Under this protocol, SAPO achieves the strongest average performance against strong APO baselines and offers an interpretable pathway for future extensions in APO."
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
Improvement: Replace monolithic prompt optimization with a four-segment decomposition (role, context, tasks, output format) and a two-stage generation pipeline.
What the improved system can do:
-
Decompose any prompt into explicit segments before optimization
-
Diagnose weak vs. strong segments using top-5/bottom-5 contrastive evidence from training data
-
Generate candidates that preserve strong segments while revising only weak ones
-
Accept updates only when validation performance improves (monotone, validation-gated search)
-
Use edit-distance tie-breaking to favor minimal, conservative edits
Improvement: For discrete metrics (e.g., ExactMatch), select top-5 positive examples and bottom-5 negative examples separately, backfilling from global rank order when needed.
Improvement: Explicitly constrain candidate generation to preserve segments labeled strong
and primarily modify segments labeled weak.
Improvement: Use one LLM with static, structured meta-prompts for segmentation, weakness analysis, and candidate generation (no dynamic prompt rewriting).
Improvement: Apply dataset-specific metrics within the optimization loop: BERTScore F1 for generation tasks, F1 for classification, ExactMatch for reasoning.
Improvement: Only accept a candidate if validation score strictly improves; otherwise keep the current prompt unchanged.
Improvement: Provide per-segment contribution analysis (role, context, tasks, output format) to identify which components drive performance.
Improvement: Use a call-level complexity proxy: SAPO requires T·Ntr + K·Nval calls per run, compared to T·K·Ntr + K·Nval for APO/StraGO.
Improvement: The method works consistently across GPT-3.5-Turbo and GPT-4o-mini without backbone-specific tuning.
Improvement: Use structured outputs (JSON-like) for diagnostic results: weak segments, strong segments, recommendations.
Summary of measurable improvements:
-
Best average score across 5 datasets on both backbones
-
+5.13% relative gain over strongest baseline on GPT-3.5-Turbo
-
+7.25% relative gain on GPT-4o-mini
-
Largest gains on GSM8K (+12.23% relative) and XSUM (+9.58% relative)
-
Consistent wins on 9 out of 10 dataset-backbone combinations
Abstract
Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation. We describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals. Using the evaluation setup across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Training Verifiers to Solve Math Word Problems
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Are Large Language Models Good Prompt Optimizers?
- Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- Modular Prompt Optimization: Optimizing Structured Prompts with Section-Local Textual Gradients
- TextGrad: Automatic "Differentiation" via Text
- Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection