MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization

arXiv:2510.16635 · cs.MA, cs.AI, cs.CL, cs.HC, cs.IR · Submitted 2025-10-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization".

Jane: Prompt optimization has become a practical way to improve Large Language Model (LLM) performance without retraining, but existing frameworks often treat evaluation as a black box,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let’s start by looking at the title and who wrote this paper—"MA-SAPO: Multi-Agent Reasoning for Score Aware Prompt Optimization." It immediately tells us that this framework uses multiple agents to focus on scores during prompt optimization.

Jane: And the authors are from Peking University, Fudan University, and the University of North Texas, which suggests a strong academic foundation behind this new approach.

Lu: The title itself points toward moving beyond simple prompting; it’s about using reasoning across different agents to make score-aware decisions during optimization. That kind of explicit reasoning structure is something I think has huge potential for how we design complex AI interactions.

Meng: It sounds like they’re trying to bridge the gap between the high-level goal of prompt optimization and the low-level execution by creating a structured pipeline instead of just letting things happen implicitly.

Lalam: Exactly, and I see this as a major step toward building more controllable AI because if we can map evaluation outcomes directly to specific edits, we gain much better control over the output quality.

The paper's summary: Tom: Now that we have the context, let’s talk about what MA-SAPO actually does according to the summary. Basically, they introduce a sequential reasoning framework where evaluation scores are used to generate targeted refinements instead of just accepting them as final results.

Jane: So, in simple terms, it means they have three main agents in the training phase: one that explains the scores, one that finds the weaknesses and trade-offs, and another that creates the actual edits we can use.

Lu: That decomposition into distinct roles—explaining, diagnosing, and synthesizing actions—is what I find fascinating; it’s like building a miniature expert team to handle prompt refinement instead of relying on one monolithic process.

Meng: The paper emphasizes that these outputs are stored as "reasoning assets," which is key because it means we aren't just throwing away the intermediate thoughts; we are saving structured knowledge for later use.

Lalam: I agree, and those assets mean that if we want to optimize a different prompt later, we can potentially retrieve those prior diagnostic insights to speed up the process significantly.

The paper's improvements: Tom: The paper outlines several key improvements they are making over existing methods. They’re addressing the black-box evaluation issue by ensuring that every refinement is directly linked back to concrete evidence from the scoring process.

Jane: It’s about shifting from implicit trial and error to an explicit pipeline where the system knows precisely what needs changing based on what it learned during evaluation.

Lu: I noticed they are specifically targeting the limitation where reasoning in current methods is often implicit; MA-SAPO aims to distill that reasoning into auditable artifacts that drive targeted edits, which is a significant architectural improvement.

Meng: From an engineering standpoint, the benefit here seems to be efficiency; by having these reusable assets, we should be able to make precise adjustments without needing a massive amount of iterative interaction every single time we test a new prompt.

Lalam: And this controllability is huge for us because it means we can track exactly how much each dimension—like helpfulness or accuracy—is improving, instead of just hoping the overall score goes up.

Conclusion: Tom: So, to wrap things up on "MA-SAPO: Multi-Agent Reasoning for Score Aware Prompt Optimization," it really boils down to taking evaluation scores and transforming them into structured reasoning assets that guide a precise refinement loop.

Jane: That means we get interpretable, auditable optimization where the system can explain exactly why a prompt got better or worse in specific areas.

Lu: I think the real implication here is that we start treating prompt tuning less like guesswork and more like a structured scientific process with defined steps for diagnosis and action.

Meng: For practical application, this framework offers a way to make precise adjustments while keeping the overall computational cost manageable compared to running full multi-agent debates every time.

Lalam: Ultimately, this paper gives us a tool that allows us to build prompt systems that are controllable and transparent, which is essential for deploying robust AI in real-world scenarios.

Tom: It’s exciting stuff, folks; MA-SAPO offers a solid blueprint for making our prompt tuning processes much more systematic and understandable.

Enhans 2Peking University 3Fudan University 4University of North Texas

cs.MA, cs.AI, cs.CL, cs.HC, cs.IR

Submitted: 2025-10-18

Updated: 2026-09-27

Importance score: 92/100

The gist: Prompt optimization has become a practical way to improve Large Language Model (LLM) performance without retraining, but existing frameworks often treat evaluation as a black box, relying solely on

Key concepts

Reasoning Assets (Ri)
These are structured outputs created during training by three agents: an Explainer, a Diagnostician, and an Action Synthesizer. They capture the 'why' behind scores—explaining the metrics, identifying error sources and trade-offs, and generating concrete edit directives. These assets serve as reusable knowledge for later optimization.
Training Phase Agents
This phase involves three specialized agents: Gexp (Metric Explainer), Gdiag (Diagnostician), and Gsyn (Action Synthesizer). They sequentially process annotated scores to build the reasoning assets. This structured approach ensures that every piece of diagnostic information is organized into a format usable by the test phase for targeted prompt refinement.
Test Phase Pipeline
This phase uses a retrieval-augmented generation pipeline. It first retrieves relevant training examples based on token similarity and then uses an Analyzer Agent to compare the new prompt against these examples, generating an improvement report. Finally, a Refiner Agent incorporates this evidence to generate the optimized prompt.
Score-Aware Optimization
Instead of blindly optimizing for one metric, MA-SAPO improves multiple quality aspects simultaneously. The framework uses structured reasoning to ensure edits are grounded in diagnostic evidence, leading to gains across several dimensions while maintaining semantic intent and efficiency.

Terminology

Summary

Prompt optimization has become a practical way to improve Large Language Model (LLM) performance without retraining, but existing frameworks often treat evaluation as a black box, relying solely on outcome scores without explaining why prompts succeed or fail. This paper introduces MA-SAPO: a new MultiAgent Reasoning for Score Aware Prompt Optimization framework that links evaluation outcomes directly to targeted refinements.

How it works

MA-SAPO is a sequential reasoning framework with two main phases: the Training Phase and the Test Phase, designed to map metric outcomes to actionable edits via structured reasoning artifacts. In the Training Phase, three agents convert annotated scores into reusable assets: (1) a Metric Explainer interprets evaluation dimensions, (2) a Diagnostician identifies error sources and trade-offs, and (3) an Action Synthesizer produces concrete edit directives. These outputs are stored as reasoning assets denoted as Ri = (Ci, Di, Ei), which form the retrieval corpus for the test phase.

Training Phase: Reasoning Asset Construction

The training phase involves three sequential agents operating on each annotated triplet τi = (pi, ri, Si). The Metric Explainer Agent generates a reasoning card Ci to explain why scores were assigned. The Diagnostician Agent extends this by producing a diagnostic summary Di that identifies: (1) the key causes of low-scoring dimensions, and (2) trade-offs across metrics. Finally, the Action Synthesizer Agent converts these insights into actionable edit directives (EDs), which are designed to be machine-parseable and directly usable by downstream agents in the test phase.

Test Phase: Retrieval and Reformulation

In the Test Phase, optimization is performed via a retrieval-augmented generation pipeline. First, a sparse lexical retriever ranks training prompts based on token overlap to retrieve the top-k most relevant prompt-response pairs along with their reasoning assets. The Analyzer Agent then compares the test prompt against these retrieved examples to generate an improvement report A, which transforms raw retrieval into structured insights by highlighting concrete weaknesses and improvement opportunities. Finally, the Refiner Agent regenerates an optimized prompt pˆ by incorporating this report, ensuring edits are grounded in diagnostic evidence and produce a final optimized prompt pˆ.

Key Components and Contributions

The framework utilizes specialized agents with distinct roles: Gexp (Metric Explainer), Gdiag (Diagnostician), Gsyn (Action Synthesizer) in training, and Gana (Analyzer) and Gref (Refiner) in testing. The core contribution is a score-aware training pipeline that distills evaluation outcomes into reusable, semi-structured reasoning assets for explanation, diagnosis, and edit directives. This results in MA-SAPO being a modular and compute-efficient framework that delivers interpretable, auditable, and controllable gains while reducing token and API-calls.

Experimental Results

Experiments on the HelpSteer1/2 benchmarks show that MA-SAPO consistently outperforms single-pass prompting, retrieval-augmented generation, and prior multi-agent methods across multiple evaluation metrics. Specifically, Table 3 demonstrates that MA-SAPO achieves superior performance compared to baselines like MARS and RAG. Furthermore, the analysis shows a consistent shift toward higher values across all five dimensions in the score distributions for both GPT-4o and Llama3-8B, indicating that MA-SAPO improves multiple quality aspects simultaneously rather than over-optimizing a single metric. The framework also maintains efficiency, achieving favorable trade-offs in both token cost and latency compared to heavier multi-agent baselines.

Human Evaluation

Qualitative analysis confirms the framework's utility through human evaluation. Annotators found that MA-SAPO outperformed the single-agent baseline in reasoning quality (H1), showing statistically significant improvements in usefulness, factual accuracy, and consistency. Additionally, evaluators rated directional consistency (H2) highly, finding that optimized prompts largely preserved the original semantic intent with a mean rating of 3.36 out of 4.0. This demonstrates that MA-SAPO improves output quality without distorting the original task intent.

Limitations and Future Work

The paper notes limitations, such as the potential for retrieval to suffer from limited recall under paraphrase or lexical mismatch, which can create a bottleneck at the retrieval stage. Additionally, the sequential pipeline is prone to cascading errors and provides limited exploration. Future work suggests improving robustness by adding verifier-based artifact validation, uncertainty-aware gated updates, and constrained rewriting, as well as introducing multi-path candidate generation with iterative verification.

Conclusion

MA-SAPO proposes a novel multi-agent framework that explicitly transforms evaluation outcomes into reusable reasoning assets and uses them to support evidence-backed refinement through the Analyzer-to-Refiner pipeline, yielding "interpretable and controllable optimization.

Improvements for AI systems

Here are specific improvements to existing AI systems based on the MA-SAPO framework, focusing on enhanced interpretability, controllability, and performance across various prompt optimization tasks:


The core improvement is shifting prompt optimization from a black-box trial-and-error process to an explicit, evidence-backed reasoning pipeline grounded in structured reasoning assets.

Here are specific ways to improve AI systems using MA-SAPO:

  1. A system can be improved by implementing a self-correcting loop where, after generating an optimized prompt, the system doesn't just check the final score but automatically executes the MA-SAPO pipeline (Analyzer and Refiner) against a set of retrieved exemplars. This ensures that every refinement is explicitly grounded in diagnostic evidence rather than heuristic guesswork.

  2. A system can be made more controllable by introducing reusable reasoning assets (the output of the Training Phase: Metric Explainer, Diagnostician, Action Synthesizer). These assets allow developers to audit and understand exactly why a prompt succeeded or failed (e.g., The helpfulness score is low because the response lacked concrete use cases). This enables precise, targeted fixes rather than broad retraining attempts.

  3. A system can achieve superior performance on complex, multi-metric tasks (like those in HelpSteer) by replacing iterative debate methods with the structured Analyzer-to-Refiner pipeline. This allows the system to simultaneously optimize for multiple quality dimensions—such as increasing helpfulness while maintaining coherence—leading to holistic improvements rather than single-metric over-optimization.

  4. For knowledge-intensive or open-ended tasks, a system can be improved by using the Test Phase's Analyzer agent to identify and enforce missing constraints or ambiguities in the prompt (as seen in Case 2). The system will then automatically generate a new prompt that explicitly defines domain boundaries, required formats, or necessary perspectives (e.g., enforcing specific structural flows like Warm-Up, Divergent Ideation, Convergent Refinement).

  5. The system can be optimized for cost and efficiency by shifting the heavy reasoning load offline during the Training Phase. This means that once a robust corpus of reasoning assets is built, optimizing new prompts only requires lightweight retrieval and a single Analyzer-to-Refiner step at test time, drastically reducing API calls and latency compared to full multi-agent debate frameworks (like MARS).

The improved AI system can perform the following specific actions:

  1. Generate highly structured, pedagogically grounded outputs by enforcing explicit formatting constraints on the LLM's response (e.g., forcing a 10-point list with required explanations instead of generic advice).

  2. Diagnose and resolve trade-offs between conflicting metrics (e.g., identifying whether a prompt is suffering from poor coherence due to excessive verbosity, or vice versa) by analyzing the interplay between reasoning cards and diagnostic summaries.

  3. Perform nuanced, criteria-driven assessments on subjective terms by decomposing them into measurable components (e.g., breaking down strongest character into physical strength, strategic skills, and magical abilities) and then optimizing the prompt to explicitly demand evaluations across all these sub-criteria.

  4. Maintain semantic stability during optimization by ensuring that when a prompt is improved for clarity or structure, the core intent of the original user query remains perfectly preserved, as verified by human evaluation metrics (H2).

  5. Scale prompt maintenance efficiently across large libraries of prompts by creating a retrievable corpus of how-to-fix instructions, allowing the system to rapidly apply proven diagnostic strategies to new inputs without requiring expensive, multi-round agent interactions for every single update.

Sources

Related papers