Maximum Likelihood Reinforcement Learning

arXiv:2602.02710 · cs.LG · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Maximum Likelihood Reinforcement Learning".

Jane: The paper was written by Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora et al. from Carnegie Mellon University and Tsinghua University and Zhejiang University and University of California, Berkeley and Impossible, Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are opening the show with a heavyweight paper called "Maximum Likelihood Reinforcement Learning".

Jane: The title alone suggests they are trying to bridge two massive worlds, Tom.

Tom: You mean the predictive side and the decision-making side?

Jane: Exactly, and the author list is incredibly impressive.

Tom: It looks like a massive collaboration between CMU and Tsinghua.

Jane: There are researchers from Zhejiang and UC Berkeley on this too.

Lu: Seeing Tsinghua and CMU working together on this suggests a very high level of theoretical rigor.

Tom: Do you think that kind of partnership is becoming the new norm for these big breakthroughs?

Lu: It certainly feels that way when the problems get this complex.

Meng: I am curious if this level of institutional coordination is what's actually needed to solve these scaling issues.

Jane: It seems like they needed a lot of brainpower to pull this off.

Meng: It definitely looks like a serious effort to move beyond the current limitations of how we train models.

Lalam: When researchers from different corners of the globe combine their expertise like this, it often changes how we view the limits of intelligence.

Tom: That is a pretty profound way to look at it, Lalam.

Jane: It really sets the stage for what they are actually proposing in the text.

Tom: We should probably get into what they are actually trying to fix.

Summary: Tom: We've introduced the team, but now we need to talk about the actual problem in "Maximum Likelihood Reinforcement Learning".

Jane: Most current reinforcement learning only cares about the average success rate.

Tom: So it's basically just trying to maximize the chance of being right on average?

Jane: Yes, and that means it doesn't put enough weight on the really difficult tasks.

Tom: You mean the ones where the model almost always fails?

Jane: Exactly, because those low-probability successes don't move the needle much in standard training.

Lu: The paper uses a Maclaurin expansion to prove that standard RL is just a very rough, first-order approximation.

Tom: That sounds like we've been using a blurry lens to look at the math.

Lu: It is quite beautiful because it shows that the true goal, which is maximum likelihood, is actually an infinite series of these successes.

Meng: So you're saying the current way we train models for math or coding is basically leaving performance on the table?

Jane: That is a good way to put it, Meng.

Meng: It feels like we are optimizing for the easy wins instead of the hard mastery.

Lalam: If we want models to truly master reasoning, we have to focus on the moments where they almost get it right.

Tom: That makes sense, but how do you actually train for that without making the math impossible?

Jane: That is where their new framework comes in.

Improvements: Tom: We are looking at how they actually implement this with "Maximum Likelihood Reinforcement Learning".

Jane: They introduced MaxRL, which uses more sampling to get a better target.

Tom: It's like trading more compute for a sharper focus, isn't it?

Jane: Yes, and they found a way to make the math work with a simple estimator.

Tom: They change how they normalize the rewards, right?

Jane: Instead of dividing by the total number of samples, they divide by the number of successful ones.

Meng: From my side, seeing that it scales better with both data and compute is what makes this a production-ready idea.

Tom: The results they show for the Qwen3 models are absolutely wild.

Lu: The twenty times efficiency gain in test-time scaling is the part that really stands out to me.

Meng: If I am reading this right, they have found a way to make the training actually more efficient as you add more samples.

Lu: It means you don't just get more stable training, you actually get a better objective.

Lalam: This efficiency could fundamentally change how we deploy intelligence in everyday tools.

Tom: It could mean we get much smarter models without needing a massive increase in energy.

Jane: We should probably wrap this up before we get too carried away.

Conclusion: Tom: We have covered a lot of ground on "Maximum Likelihood Reinforcement Learning".

Jane: It really feels like a fundamental shift in how we approach correctness-based training.

Tom: We've seen how they move from simple averages to a much more precise mathematical target.

Jane: And the scaling benefits they found for reasoning tasks are hard to ignore.

Lu: I think this marks the beginning of a much more rigorous era for reinforcement learning.

Meng: I am looking forward to seeing how this changes the actual engineering pipelines for large models.

Lalam: This approach will allow AI to bridge the gap between simple imitation and true logical mastery.

Tom: Thanks to everyone for joining us to break this down.

Jane: We'll see you next time for the next big paper.

Tom: Goodbye for now!

Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Yiding Jiang, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette

Carnegie Mellon University · Tsinghua University · Zhejiang University · University of California, Berkeley · Impossible, Inc.

cs.LG

Submitted: 2026-08-21

Updated: 2026-08-24

Project page: https://zanette-labs.github.io/MaxRL

Importance score: 80/100

The gist: The paper introduces "Maximum Likelihood Reinforcement Learning (MaxRL), a sampling-based framework to approximate maximum likelihood using reinforcement learning techniques." The authors observe

Key concepts

Standard Reinforcement Learning Limitations
Current reinforcement learning often focuses on average success rates, which fails to prioritize difficult tasks where models frequently fail. The paper demonstrates that standard RL is actually a rough, first-order approximation of the true mathematical goal, which is maximum likelihood.
MaxRL
MaxRL is a framework that uses increased sampling to achieve a sharper focus during training. It improves efficiency by changing how rewards are normalized, dividing them by the number of successful samples rather than the total number of samples, which helps the model scale better.
Test-time Scaling
This involves how much a model's performance improves as more compute or samples are used during testing. The MaxRL framework achieved a twenty times efficiency gain in test-time scaling for Qwen3 models, helping models bridge the gap between simple imitation and true logical mastery.

Terminology

Summary

The paper introduces Maximum Likelihood Reinforcement Learning (MaxRL), a sampling-based framework to approximate maximum likelihood using reinforcement learning techniques. The authors observe that in correctness-based settings—such as navigation, code generation, and mathematical problem solving—reinforcement learning does not maximize [the implicit likelihood], and instead optimizes only a lower-order approximation. They formalize this by showing that standard reinforcement learning optimizes only the first-order approximation of the maximum likelihood objective.

To bridge this gap, the authors define a compute-indexed family of sample-based objectives that interpolate between standard reinforcement learning and exact maximum likelihood as additional sampling compute is allocated. This framework is built upon the observation that the maximum likelihood objective admits the Maclaurin expansion in terms of failure events: J ML(x) = p = -sum k=1 infinity(1-p) k over k, which implies the population-level gradient identity grad theta J ML(x) = sum k=1 infinity 1 over k grad theta pass@k(x). MaxRL addresses the difficulty of estimating this infinite sum by using a truncated version: J MaxRL(T)(x):= -sum k=1 T(1-p) k over k, where the truncation level T controls the order of correctness events that contribute to learning.

The paper provides theoretical grounding for the gradient estimation, stating that the gradient of the maximum likelihood objective admits the following conditional expectation representation: grad theta J ML(x) = E[grad theta m theta(zx) f(z) = y*(x)]. The authors propose an empirical estimator that averages score functions only over successful trajectories and prove that the estimator N(x) is an unbiased estimator for the MaxRL gradient of order T=N. A fundamental distinction is made between MaxRL and standard REINFORCE: REINFORCE reduces variance of a fixed objective (pass@1), while MaxRL increases the approximation order to maximum likelihood.

Empirical evaluations across various settings demonstrate that MaxRL Pareto-dominates existing methods in all models and tasks we tested, achieving up to 20× test-time scaling efficiency gains compared to its GRPO-trained counterpart. The findings include:

  • Controlled Settings: In image classification, MaxRL improves consistently and closely tracks exact maximum likelihood as compute increases.

  • Infinite Data Regime: In maze-navigation tasks, MaxRL scales with compute far more effectively than competing frameworks when large amounts of unique training data is available.

  • Data-Scarce Regime: On the GSM8K dataset, MaxRL is more resistant to overfitting... demonstrating less pass@k degradation (overfitting) and converging to a higher average performance.

  • Large-Scale Reasoning: When training Qwen3-1.7B and Qwen3-4B models on mathematical reasoning, MaxRL Pareto-dominates GRPO, shows little to no diversity degradation with respect to the base model, and leads to strong (up to 20×) test-time scaling efficiency gains.

Furthermore, the authors note that MaxRL exhibits characteristically different optimization dynamics, specifically that it produces stronger gradients on harder prompts and leads to a larger fraction of prompts with at least one correct rollout during training.

Improvements for AI systems

1. Implementation of MaxRL Gradient Estimators in Correctness-Based Training

  • Improvement: Replace standard policy-gradient advantage functions (such as those used in REINFORCE, RLOO, or GRPO) with the MaxRL estimator: N(x) = 1 over K sum i=1 K grad theta m theta(z i x), where K is the number of successful rollouts out of N total samples.

  • Capability: The AI system will transition from optimizing a first-order approximation (expected reward) to a higher-order approximation of the maximum likelihood objective. This enables the model to prioritize hard inputs—those with low success rates—by generating higher gradient norms for difficult prompts, preventing the model from stalling on complex reasoning tasks that standard RL ignores.

2. Compute-Indexed Training Curriculum

  • Improvement: Integrate a training schedule where the rollout budget N is treated as a hyperparameter that scales with allocated compute, effectively performing a Maclaurin expansion of the likelihood objective during the training process.

  • Capability: The AI system will exhibit superior scaling laws. Unlike standard RL, where increasing compute only reduces gradient variance, increasing compute in a MaxRL framework improves the actual optimization objective, allowing the model to converge toward exact maximum likelihood as more sampling compute is utilized.

3. Optimized Test-Time Scaling (Verifier-in-the-Loop)

  • Improvement: Deploy MaxRL-trained models within search-based inference frameworks (e.g., Best-of-N or Majority Voting) paired with a perfect or high-quality verifier.

  • Capability: The AI system will achieve up to 20× gains in test-time scaling efficiency. Because MaxRL prevents distribution sharpening (the collapse of output diversity), the model maintains a high pass@k, allowing it to solve complex mathematical or coding problems with significantly fewer total inferences compared to GRPO-trained models.

4. Diversity-Preserving Post-Training for Data-Scarce Regimes

  • Improvement: Utilize MaxRL for reinforcement learning on fixed, high-quality reasoning datasets (e.g., GSM8K) instead of standard RLVR (Reinforcement Learning from Verifiable Rewards).

  • Capability: The AI system will be highly resistant to overfitting and mode collapse. While standard RL models tend to lose reasoning diversity (dropping in pass@k) as they over-optimize on a fixed dataset, a MaxRL-trained system will sustain higher entropy and a broader coverage of valid reasoning paths, ensuring robust generalization to unseen prompts.

Sources

Related papers