P 2O: Joint Policy and Prompt Optimization

summary

Video file (mp4)

The gist

Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals, and this paper introduces Joint Policy and

In short

The paper introduces Joint Policy and Prompt Optimization (P2O) to fix advantage collapse in Reinforcement Learning with Verifiable Rewards (RLVR). P2O alternates between updating model parameters and evolving discrete prompts to handle hard samples. This joint approach prevents the model from losing learning signals on difficult problems, leading to significant performance gains over standard methods.

Key concepts

Joint Optimization
This framework optimizes both the continuous policy parameters (theta) of the AI model and a set of discrete prompt templates (Z) simultaneously. Instead of just updating the model or just changing prompts, P2O finds the best combination of both to maximize performance on hard samples.
Context Distillation
This technique is used during policy updates to transfer reasoning skills from a prompt directly into the model's parameters. It forces the AI to internalize prompt-guided steps into its core knowledge, making it less reliant on prompts during actual problem-solving.
Evolutionary Prompt Optimization (GEPA)
This phase uses an LLM called a 'Reflection LLM' to analyze why current prompts fail on hard samples. It then generates new, targeted prompt mutations to fix those specific errors. This iterative process ensures the prompts evolve to cover more reasoning patterns.
Unified Policy Gradient
The policy gradient is modified to combine standard learning from easy samples with guided learning from hard samples using different contexts (x and x_tilde). This allows the model to learn effectively from both simple and complex examples in a single update step.

Terminology used across episodes

This episode discusses

The paper

P 2O: Joint Policy and Prompt Optimization · Read on arXiv

Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · National Computer Network Emergency Response Technical Team/Coordination Center of China

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "P 2O: Joint Policy and Prompt Optimization".

Jane: Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We started by looking at the title "P 2O: Joint Policy and Prompt Optimization," which immediately tells us that this work is about combining two different optimization strategies to improve model performance. The authors are Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu, Hongyu Lin, Min He, and Xianpei Han.

Jane: Those authors seem like a solid group of researchers coming from the Chinese Academy of Sciences and related institutes; they've clearly got deep expertise in the underlying machine learning and information processing areas.

Lu: Their background seems very well-suited for this work because it touches on both continuous policy updates and discrete prompt evolution, which is a complex area to manage effectively.

Meng: I wonder if their focus on joint optimization means they're trying to solve a problem that single methods just can't handle, which is usually the case when we look at hard samples in RL.

Lalam: It suggests that the next step for AI isn't just making the model bigger or training it longer, but developing more sophisticated ways to guide its learning process directly through prompt design and parameter adjustment simultaneously.

The paper's summary: Tom: So, the core of "P 2O: Joint Policy and Prompt Optimization" is introducing this joint optimization framework where they maximize a joint objective over both the policy parameters theta and the discrete prompt templates Z. They break it down into two main phases.

Jane: That sounds like they are treating the model's internal settings and its external instructions as things that need to be tuned together, which is a very holistic approach to problem-solving.

Lu: The first phase involves Policy Optimization with Context Distillation, where the model learns from prompt-guided trajectories but updates its parameters using the original input x, forcing it to learn those newly discovered reasoning steps internally rather than just relying on the prompt during inference.

Meng: That context distillation part is interesting; it’s about embedding the learned reasoning directly into the policy parameters theta so that we don't need to rely on that specific prompt every single time we run an input.

Lalam: That internalizing of capabilities sounds like it moves us closer to a more robust and less brittle AI system, one where the knowledge is truly part of its core structure.

Tom: Then they move into the second phase, Evolutionary Prompt Optimization, where they use an algorithm called GEPA to evolve a new set of prompts targeting those remaining hard samples identified in the first phase.

Jane: So, it’s a cycle: you optimize parameters based on what you can do with the current prompts, and then you evolve those prompts specifically to tackle the hardest problems left over.

Lu: It’s about systematically finding and internalizing successful reasoning paths through this alternating process, moving from exploration guided by simple samples to targeted evolution for the most difficult cases.

The paper's improvements: Tom: The paper highlights that P 2O improves upon standard approaches like Group Relative Policy Optimization, which they call GRPO; they show it breaks the rollout-scaling ceiling that vanilla GRPO hits. They claim a robust nine percent improvement over those runs on DeepMath-5K under compute-equivalent conditions.

Jane: A nine percent improvement is substantial when you think about how much effort goes into training these large models, so showing that gain without just scaling up the compute is quite impressive.

Meng: If it’s improving performance by leveraging prompts to reach high-reward regions inaccessible through standard exploration, that suggests we can tackle problems that were previously considered too difficult for the model to solve reliably.

Lalam: And they also show that this diversity in reasoning space coverage, achieved by evolving templates, leads to pronounced gains on benchmarks like AIME24 and AIME25, which is great proof of concept for real-world reasoning.

Lu: The results suggest that these evolved templates don't just help with a single lucky guess; they robustly enhance the model’s solution space, boosting both the deterministic success rate and the broader exploration coverage needed for effective distillation.

Tom: They also point out that excluding context distillation causes a severe drop in accuracy, which confirms that optimizing on the prompt rather than on x induces a dependency effect we need to be careful about.

Jane: It’s important to note what they say about limitations; they mention that the optimal reflection source for GEPA is task-dependent, meaning it changes depending on what kind of problem you're looking at.

Meng: That dependency in the evolution phase means we can't just use one universal prompt generator; we have to tailor the prompting mechanism to specific failure modes, which makes sense for practical deployment.

Conclusion: Tom: To wrap up "P 2O: Joint Policy and Prompt Optimization," the paper confirms that this iterative self-improvement process by decoupling discrete semantic search and continuous parameter updates is a robust way to achieve autonomous self-improvement in large language models.

Jane: It really shows that joint optimization is a viable pathway for making AI systems capable of tackling problems they couldn't solve on their own before, especially when facing those hard samples.

Lu: The whole idea of alternating between policy updates and prompt evolution proves that we can systematically discover and internalize the correct reasoning trajectories needed for intractable problems through this joint optimization method.

Meng: From an engineer’s view, this framework suggests a pathway to building AI agents that don't just rely on brute-force sampling but instead have an intelligent, adaptive loop for refining their own instructional strategies.

Lalam: Ultimately, P 2O establishes a cycle where the AI learns from its successes by evolving its instructions and updating its core weights based on those successful interactions with hard problems.

Tom: What we've seen here is a very structured approach to handling the limits of current RL methods when dealing with complex reasoning tasks, and it sets a new direction for how we can build more capable models.

More episodes

← Home