P 2O: Joint Policy and Prompt Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "P 2O: Joint Policy and Prompt Optimization".
Jane: Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We started by looking at the title "P 2O: Joint Policy and Prompt Optimization," which immediately tells us that this work is about combining two different optimization strategies to improve model performance. The authors are Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu, Hongyu Lin, Min He, and Xianpei Han.
Jane: Those authors seem like a solid group of researchers coming from the Chinese Academy of Sciences and related institutes; they've clearly got deep expertise in the underlying machine learning and information processing areas.
Lu: Their background seems very well-suited for this work because it touches on both continuous policy updates and discrete prompt evolution, which is a complex area to manage effectively.
Meng: I wonder if their focus on joint optimization means they're trying to solve a problem that single methods just can't handle, which is usually the case when we look at hard samples in RL.
Lalam: It suggests that the next step for AI isn't just making the model bigger or training it longer, but developing more sophisticated ways to guide its learning process directly through prompt design and parameter adjustment simultaneously.
The paper's summary: Tom: So, the core of "P 2O: Joint Policy and Prompt Optimization" is introducing this joint optimization framework where they maximize a joint objective over both the policy parameters theta and the discrete prompt templates Z. They break it down into two main phases.
Jane: That sounds like they are treating the model's internal settings and its external instructions as things that need to be tuned together, which is a very holistic approach to problem-solving.
Lu: The first phase involves Policy Optimization with Context Distillation, where the model learns from prompt-guided trajectories but updates its parameters using the original input x, forcing it to learn those newly discovered reasoning steps internally rather than just relying on the prompt during inference.
Meng: That context distillation part is interesting; it’s about embedding the learned reasoning directly into the policy parameters theta so that we don't need to rely on that specific prompt every single time we run an input.
Lalam: That internalizing of capabilities sounds like it moves us closer to a more robust and less brittle AI system, one where the knowledge is truly part of its core structure.
Tom: Then they move into the second phase, Evolutionary Prompt Optimization, where they use an algorithm called GEPA to evolve a new set of prompts targeting those remaining hard samples identified in the first phase.
Jane: So, it’s a cycle: you optimize parameters based on what you can do with the current prompts, and then you evolve those prompts specifically to tackle the hardest problems left over.
Lu: It’s about systematically finding and internalizing successful reasoning paths through this alternating process, moving from exploration guided by simple samples to targeted evolution for the most difficult cases.
The paper's improvements: Tom: The paper highlights that P 2O improves upon standard approaches like Group Relative Policy Optimization, which they call GRPO; they show it breaks the rollout-scaling ceiling that vanilla GRPO hits. They claim a robust nine percent improvement over those runs on DeepMath-5K under compute-equivalent conditions.
Jane: A nine percent improvement is substantial when you think about how much effort goes into training these large models, so showing that gain without just scaling up the compute is quite impressive.
Meng: If it’s improving performance by leveraging prompts to reach high-reward regions inaccessible through standard exploration, that suggests we can tackle problems that were previously considered too difficult for the model to solve reliably.
Lalam: And they also show that this diversity in reasoning space coverage, achieved by evolving templates, leads to pronounced gains on benchmarks like AIME24 and AIME25, which is great proof of concept for real-world reasoning.
Lu: The results suggest that these evolved templates don't just help with a single lucky guess; they robustly enhance the model’s solution space, boosting both the deterministic success rate and the broader exploration coverage needed for effective distillation.
Tom: They also point out that excluding context distillation causes a severe drop in accuracy, which confirms that optimizing on the prompt rather than on x induces a dependency effect we need to be careful about.
Jane: It’s important to note what they say about limitations; they mention that the optimal reflection source for GEPA is task-dependent, meaning it changes depending on what kind of problem you're looking at.
Meng: That dependency in the evolution phase means we can't just use one universal prompt generator; we have to tailor the prompting mechanism to specific failure modes, which makes sense for practical deployment.
Conclusion: Tom: To wrap up "P 2O: Joint Policy and Prompt Optimization," the paper confirms that this iterative self-improvement process by decoupling discrete semantic search and continuous parameter updates is a robust way to achieve autonomous self-improvement in large language models.
Jane: It really shows that joint optimization is a viable pathway for making AI systems capable of tackling problems they couldn't solve on their own before, especially when facing those hard samples.
Lu: The whole idea of alternating between policy updates and prompt evolution proves that we can systematically discover and internalize the correct reasoning trajectories needed for intractable problems through this joint optimization method.
Meng: From an engineer’s view, this framework suggests a pathway to building AI agents that don't just rely on brute-force sampling but instead have an intelligent, adaptive loop for refining their own instructional strategies.
Lalam: Ultimately, P 2O establishes a cycle where the AI learns from its successes by evolving its instructions and updating its core weights based on those successful interactions with hard problems.
Tom: What we've seen here is a very structured approach to handling the limits of current RL methods when dealing with complex reasoning tasks, and it sets a new direction for how we can build more capable models.
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · National Computer Network Emergency Response Technical Team/Coordination Center of China
cs.LG, cs.AI
Submitted: 2026-03-23
Updated: 2026-09-28
Code: https://github.com/QwenLM/Qwen2.5-Math
Importance score: 90/100
The gist: Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals, and this paper introduces Joint Policy and
Key concepts
- Joint Optimization
- This framework optimizes both the continuous policy parameters (theta) of the AI model and a set of discrete prompt templates (Z) simultaneously. Instead of just updating the model or just changing prompts, P2O finds the best combination of both to maximize performance on hard samples.
- Context Distillation
- This technique is used during policy updates to transfer reasoning skills from a prompt directly into the model's parameters. It forces the AI to internalize prompt-guided steps into its core knowledge, making it less reliant on prompts during actual problem-solving.
- Evolutionary Prompt Optimization (GEPA)
- This phase uses an LLM called a 'Reflection LLM' to analyze why current prompts fail on hard samples. It then generates new, targeted prompt mutations to fix those specific errors. This iterative process ensures the prompts evolve to cover more reasoning patterns.
- Unified Policy Gradient
- The policy gradient is modified to combine standard learning from easy samples with guided learning from hard samples using different contexts (x and x_tilde). This allows the model to learn effectively from both simple and complex examples in a single update step.
Terminology
Summary
Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals, and this paper introduces Joint Policy and Prompt Optimization (P2O) to mitigate this collapse by alternating continuous policy updates with discrete prompt evolution.
How it works
-
P2O alternates between two phases: (1) Policy Optimization with Context Distillation, where the model learns from prompt-guided trajectories but updates its parameters using the original input, forcing it to internalize the newly discovered reasoning steps. (2) Evolutionary Prompt Optimization, where GEPA is applied to evolve new prompts targeting remaining hard samples.
-
The framework reformulates the RL objective as a joint optimization over policy parameters θ and discrete prompt templates Z:
max θ,Z Jjoint = X x∈D·Dhard max z∈Z Ey∼πθ(·T (x,z))[r(x, y)] (2)
- In Phase 1, the model identifies hard samples D(t+1) hard by filtering instances whose empirical success rate falls below a threshold τ:
D(t+1) hard = (xi ∈ D(t) hard and X K k=1 r(xi, yik) < τ) (3)
- For a hard sample x ∈ D(t) hard, the policy gradient is unified by combining standard exploration on simple samples with prompt-guided exploration on hard samples:
∇θJjoint ≈ X x /∈D(t) hard Ey∼πθ(·x) A(x, y)∇θ log πθ(yx) + X x∈D(t) hard Ey˜∼πθ(·x˜) A(x, y˜)∇θ log πθ(˜yx) (4
Key Mechanisms
(1)
Context Distillation is employed in Phase 1 to transfer the reasoning capabilities triggered by a prompt z directly into the parameters of πθ for the original input x. This prevents dependency on inference-time prompting. The unified policy gradient decouples the rollout context (x˜) from the gradient context (x), effectively distilling
prompt-elicited reasoning into the policy parameters θ.
(2)
The Evolutionary Prompt Optimization phase utilizes GEPA to evolve the template set Z(t+1) to address D(t+1) hard. GEPA employs a “Reflection LLM” to analyze error traces from current prompt performance and generate targeted semantic mutations that address specific failure modes. This process is guided by a meta-prompt:
z′ ∼ πinit (· promptmeta, z, F) (5)
(3)
GEPA maintains a population of diverse prompts using a Pareto-based selection mechanism to ensure diversity across failure modes. After evolution, a greedy set cover strategy is applied via Algorithm 2 to select a minimal Pareto set Zcovered and assign sample-specific templates to each instance in Dhard for the next training epoch.
Empirical Findings
(1)
P2O consistently outperforms standard GRPO and its variants by breaking the rollout-scaling ceiling of vanilla GRPO, achieving a robust 9% improvement over GRPO runs
on DeepMath-5K under compute-equivalent conditions.
(2)
P2O substantially surpasses single-turn reflection baselines, which fall below standard GRPO performance, proving that GEPA’s iterative evolutionary process is necessary to reliably discover and internalize correct reasoning trajectories. The optimal reflection source is taskdependent: Teacher-Reflection dominates on DeepScaler-5K, while P2OSelf-Ref leads on DeepMath-5K.
(3)
Ablation studies confirm the necessity of context distillation; excluding it causes a severe drop in accuracy, indicating that optimizing the model conditioned on the augmented prompt rather than on x induces a “dependency” effect. Furthermore, diversity-driven approach (P2O) outperforms single-template baseline by achieving pronounced gains on AIME24 (+2.5%) and AIME25 (+4.8%)
by broadening reasoning space coverage.
(4)
The framework demonstrates that the evolved templates do not merely facilitate a single lucky guess; rather, they robustly enhance the model’s solution space, boosting both the deterministic success rate and the broader exploration coverage required for effective distillation.
The case study on geometric reasoning illustrates how an optimized prompt injects specific domain knowledge—the concept of a regular tetrahedron—to steer reasoning toward the global optimum.
Conclusion
P2O establishes an iterative self-improvement process by decoupling discrete semantic search and continuous parameter updates, confirming that joint optimization is a robust pathway for autonomous self-improvement in LLMs.
Improvements for AI systems
Here are the specific improvements that an AI system, leveraging the Joint Policy and Prompt Optimization (P2O) framework, can achieve:
-
The AI system gains robust performance on
hard samples
where standard Reinforcement Learning with Verifiable Rewards (RLVR) fails due to advantage collapse. This means it will no longer suffer from a loss of learning signals when encountering complex or rare reasoning problems. -
It can achieve up to a 9.5% performance improvement over standard GRPO and surpass baselines by doubling rollout budgets, effectively scaling its problem-solving capability without simply increasing computational cost linearly.
-
The system will demonstrate strong out-of-distribution generalization (OOD generalization) because the P2O framework learns reasoning patterns directly into its internal parameters via context distillation, eliminating the need for costly inference-time prompting during deployment.
-
It can solve complex, multi-step mathematical and logical reasoning benchmarks (like AIME24/AIME25) with significantly higher accuracy (e.g., achieving 64.2% on DeepMath-5K), surpassing models that rely solely on standard policy optimization or single-turn reflection strategies (which score around 50%).
-
The system exhibits superior structural exploration: it can move beyond local optima by dynamically evolving prompt templates using the GEPA algorithm, which systematically discovers and internalizes the correct reasoning paths required for intractable problems.
-
The AI system's performance is resilient to variations in expert guidance; it can outperform teacher-dependent variants (like P2OTeacher-Ref) on certain datasets by learning self-generated refinements, indicating a more autonomous self-improvement cycle.
-
It can be fine-tuned for specific domains by leveraging the GEPA process, which allows the model to adapt its
latent instruction space
to address domain-specific failure modes (e.g., geometric packing vs. algebraic manipulation).
Sources
- OpenAI o1 System Card
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Outcome-based Exploration for LLM Reasoning
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Llama-Nemotron: Efficient Reasoning Models
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Learning by Distilling Context
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Qwen3 Technical Report
- Kimi K2: Open Agentic Intelligence
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
- GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
- ExGRPO: Learning to Reason from Experience
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks