MDP Planning as Policy Inference

arXiv:2602.17375 · cs.LG · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MDP Planning as Policy Inference".

Jane: The paper was written by David Tolpin from Offtopia.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Core Idea: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and with me is Jane, and we are looking at a paper that just landed on arXiv called "MDP Planning as Policy Inference." Jane, I have to say, the title alone got me excited because it's taking something classic and flipping it on its head.

Jane: It really does, Tom. And the core idea is deceptively simple. Instead of trying to find the single best action or the single best policy through the usual reinforcement learning tricks, this paper says, let's treat the policy itself as a hidden variable and do Bayesian inference over it. So you're not just asking "what should I do?" but "which whole strategy is most likely to be the right one?"

Tom: Right, and that's a big shift. Normally, in MDP planning, you're optimizing an objective, maximizing expected return. Here, they assign each policy an unnormalized probability that goes up as its expected return goes up. So the most likely policies are the ones that get the highest rewards, but you get a whole distribution over policies instead of just one winner.

Jane: And that distribution is the magic. It gives you uncertainty. If two different strategies have nearly the same expected return, both of them will have high probability in the posterior. That means when you act, you can sample from that posterior, and you get a stochastic controller that naturally reflects how confident you are about what the optimal behavior is.

Tom: So it's not just about finding the answer, it's about knowing how sure you are of the answer. And the way they act is essentially Thompson sampling, which is a classic idea, but applied here at the policy level, not just the action level.

Jane: Exactly. And I love that they're explicit about this. They say the stochasticity in the final behavior is not entropy regularization, which is what you get in a lot of soft reinforcement learning methods. It's genuine uncertainty over which deterministic policy is the right one.

Tom: And that's a crucial distinction, because entropy regularization often forces randomness even when you're certain. Here, if the posterior is concentrated, you get a nearly deterministic policy. If it's diffuse, you get randomness. The uncertainty drives the behavior, not a fixed coefficient.

Jane: So the title really captures it. "MDP Planning as Policy Inference" means you take a planning problem, you turn it into an inference problem, and you get a richer object out of it. And I'm curious, Tom, how do they actually make this computationally tractable? Because inference over policies sounds expensive.

Tom: That's exactly what we're going to dig into next, because they use a variational sequential Monte Carlo approach, and they had to make some clever modifications to make it work. Stick around.

Methodology and the Clever Modifications: Tom: So Jane, we've got the big picture, but the paper gets really interesting in the details. They use variational sequential Monte Carlo, or VSMC, to actually approximate this posterior over policies. But they had to adapt it in two specific ways, and those adaptations are the heart of the method.

Jane: Right, and the first one is about consistency. Since they're inferring deterministic policies, if a particle visits the same state twice, it has to take the same action both times. So they memoize the action for each state on first visit and reuse it on revisits. That sounds straightforward, but it's a real departure from standard SMC, where you'd sample a fresh action every time.

Tom: And the second modification is even more subtle. They couple the transition randomness across all particles within a sweep. So if two particles are in the same state and take the same action on the same visit count, they get forced to the same next state. That way, the weights of the particles reflect differences in policy, not differences in luck from the environment's stochasticity.

Jane: That's a really elegant trick. It's like running a controlled experiment. You hold the environment fixed, and then the only thing that differs between particles is the policy they're following. So when you compute the importance weights, you're isolating the effect of the policy choice.

Tom: And they prove in Theorem one that the gradient estimator they use is unbiased. So even though they're doing this coupled sampling and using a score function estimator, they show that optimizing their surrogate objective is a proper stochastic gradient ascent on the expected log evidence.

Jane: And I think the key insight for me is that they're treating the single-episode return as a noisy Monte Carlo estimate of the policy's log probability. So the randomness from the environment is just noise in the objective, and the inference machinery has to average over that noise.

Tom: Right, and that's why the coupling is so important. Without it, the noise would dominate the signal, and the particle weights would be mostly random. With it, the weights actually tell you something about which policies are better.

Jane: So this is a real algorithmic contribution, not just a conceptual one. They're not just saying "let's do Bayesian inference over policies," they're showing exactly how to make that inference work in practice, with a specific algorithm and a proof that it's doing the right thing.

Tom: And I have to say, the ablation studies in the paper really sell it. They show that if you drop the deterministic policy enforcement, you get a higher-entropy, mushier policy distribution. And if you drop the shared dynamics, the agent becomes overly cautious or overly risky in the grid world. So both modifications are load-bearing.

Jane: So the method is sound, but I'm dying to know how it actually performs against the standard baselines. I mean, we've got the algorithm, but does it win?

Tom: That's the next segment. We're going to look at the experiments, and honestly, the results are a mixed bag in the most interesting way. Let's get into it.

Experiments and Comparisons: Tom: So Jane, we've got the algorithm, and now we need to see it in action. The paper runs experiments on grid worlds, Blackjack, Triangle Tireworld, and Academic Advising, comparing against discrete Soft Actor-Critic, or SAC.

Jane: And the grid world results are really illustrative. They show that VSMC and SAC produce different policies, even when the return distributions are similar. In particular, SAC tends to push actions toward grid boundaries, because that increases entropy, but it doesn't actually help reach the goal. VSMC penalizes those actions because a deterministic policy that walks into a wall can only escape due to environment stochasticity.

Tom: That's a really concrete example of the philosophical difference we talked about earlier. SAC is optimizing for entropy, so it likes actions that keep options open. VSMC is optimizing for expected return under a deterministic policy, so it avoids actions that rely on luck.

Jane: And then Blackjack is fascinating because there's a known optimal policy. They compute it by value iteration. And what they find is that VSMC with default settings gets a higher expected reward than SAC with its default entropy weight. To get SAC to match VSMC, they have to drop the entropy weight from one to zero point one. And even then, VSMC has a lower draw probability than both the optimal policy and SAC.

Tom: So VSMC is essentially playing a more aggressive game. It's less likely to draw, which means it's more likely to win or lose outright. And that's a direct consequence of the posterior concentrating on policies that maximize expected return, without the entropy term pushing toward safer, more mixed behavior.

Jane: But then we get to Triangle Tireworld, and this is where the paper gets really honest. With the original rewards, VSMC performs poorly. The return gap between "fast but risky" and "safe but slow" is huge, so the posterior becomes extremely peaked, and the algorithm essentially commits to one behavior with high variance. But when they scale the rewards down by a factor of five, the posterior becomes less concentrated, and VSMC matches SAC.

Tom: That's a real limitation, and they own it. Classical MDP planning is invariant to affine reward scaling, but Bayesian inference is not, because the scale controls how peaked the posterior is. So the method works best when the reward scale meaningfully encodes the strength of preferences, not just the ranking of policies.

Jane: And then Academic Advising, which is a big combinatorial problem with long horizons, shows that both methods struggle on harder instances, but VSMC has heavier-tailed return distributions. So it's finding policies that sometimes do great and sometimes do terribly, while SAC is more consistent but less likely to hit the high end.

Tom: So the empirical picture is nuanced. VSMC isn't uniformly better, but it's different in ways that matter. It's more aggressive in Blackjack, more sensible in grid worlds, and more fragile in Triangle Tireworld.

Jane: And that fragility is really important for anyone who wants to use this in practice. You can't just set the reward scale arbitrarily and expect good results. You have to think about what the scale means for your uncertainty.

Tom: So let's bring in our guests. Lu, Meng, Lalam, what do you make of these results? Lu, you're the researcher, what's the big picture here?

Lu: I think the big picture is that this gives us a principled way to talk about uncertainty in planning. Instead of just saying "here's the optimal policy," you can say "here's a distribution over policies, and the spread tells you how much we don't know." That's a huge deal for safety-critical applications where you need to know when you're uncertain.

Meng: But from an engineering standpoint, I'm worried about the computational cost. VSMC with ten particles for fifty thousand iterations, that's a lot of environment calls. And the memoization and coupling require careful bookkeeping. Is this going to scale to real-world problems with continuous state spaces?

Tom: That's a fair question, and the paper addresses it briefly. They say the semantics don't depend on discreteness, and you can use hashable state abstractions or keyed random streams for continuous domains. But they don't actually show it working, so it's an open question.

Lalam: I think the most impactful vision here is that this could change how we build AI systems that need to explain their decisions. If you have a posterior over policies, you can say "we're seventy percent sure the optimal behavior is this, and thirty percent sure it's that." That kind of uncertainty communication is essential for building trust with humans.

Jane: That's a beautiful point, Lalam. And it connects back to the Thompson sampling interpretation. The agent is essentially saying "I'm not sure which strategy is best, so I'm going to randomize according to my beliefs." That's a very human way to make decisions.

Tom: So we've got a method that's principled, has some real differences from standard approaches, and opens up new questions about uncertainty and scaling. I think that's a great place to wrap up. Let's do a quick summary.

Conclusion: Tom: So, Jane, we've spent the whole show on "MDP Planning as Policy Inference," and I think we've covered a lot of ground. Let's pull it together.

Jane: Absolutely. The paper takes the classic problem of MDP planning and recasts it as Bayesian inference over policies. Each policy gets a probability proportional to its expected return, and the resulting posterior gives you both the optimal behavior and a measure of how uncertain you are about it.

Tom: And the algorithm, variational sequential Monte Carlo with those two key modifications, policy consistency and coupled dynamics, makes this tractable. The experiments show real differences from Soft Actor-Critic, with VSMC being more aggressive in Blackjack, more sensible in grid worlds, and more fragile in Triangle Tireworld.

Jane: The Triangle Tireworld result is a good reminder that this isn't a silver bullet. The reward scale matters, and you have to think about what it means for your uncertainty. But when it works, you get a richer object than just a single policy.

Tom: And the implications are big. This gives us a principled way to talk about uncertainty in planning, which is crucial for safety-critical systems and for building trust with humans. It's not just about finding the answer, it's about knowing how sure you are.

Jane: And with that, we're going to say goodbye to this paper. It's a thought-provoking piece that opens up more questions than it answers, and that's exactly what we want from a good arXiv paper.

Tom: Thanks for listening, everyone. We'll be back soon with the next paper. Until then, keep exploring.

David Tolpin

Offtopia

cs.LG

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 18 pages, many figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper "MDP Planning as Policy Inference" by David Tolpin formulates episodic Markov decision process (MDP) planning as Bayesian inference over policies.

Key concepts

MDP Planning as Policy Inference
This approach recasts the classic problem of Markov Decision Process planning into a Bayesian inference problem. Policies are treated as hidden variables, and the goal is to find a distribution over policies rather than just one optimal action.
Variational Sequential Monte Carlo (VSMC)
A computational method used in the paper to approximate the complex posterior distribution over policies. It requires two key modifications: enforcing policy consistency and coupling transition randomness across particles.
Policy Inference
The core idea is to assign an unnormalized probability to every possible policy based on its expected return. This results in a posterior distribution that quantifies not only the best strategy but also the uncertainty about it.
Soft Actor-Critic (SAC)
A standard reinforcement learning method used as a baseline. SAC optimizes for entropy, which encourages randomness and keeping options open, often leading to different behaviors than those derived from maximizing expected return.

Terminology

Summary

The paper MDP Planning as Policy Inference by David Tolpin formulates episodic Markov decision process (MDP) planning as Bayesian inference over policies. The core idea is stated in the abstract: Each policy is assigned an unnormalized density monotone in its expected return, so posterior modes recover return-optimal solutions while posterior dispersion quantifies uncertainty over optimal behavior.

The paper's stated contributions are: (1) "A formulation of episodic MDP planning as Bayesian inference over policies that preserves the classical expected-return optimality criterion and yields an optimal stochastic policy under preference uncertainty via posterior predictive sampling; (2) An adaptation of VSMC for inference over deterministic policies in discrete MDPs with stochastic transitions, including policy consistency under revisitation and coupled transition randomness across particles; and (3) An empirical evaluation of the induced stochastic control policy obtained by posterior predictive (Thompson-style) action sampling, and a comparison to discrete Soft Actor-Critic across diverse discrete benchmarks."

The probabilistic model assigns to each policy an unnormalized probability of optimality monotone in expected return. Specifically, we define the unnormalized log probability of a policy as the expected return obtained by the agent following the policy in an episode, given by log p̃(π) = E τπ [Σ t=1 H R(s t, a t, s t+1)], where a t = π(s t). The paper notes that neither actions nor states are treated as Bayesian random variables for which a posterior is sought — they are generative rather than inferential sources of randomness. The log probability is available only through noisy Monte Carlo evaluations by computing the return of a single episode.

For inference, the paper adapts variational sequential Monte Carlo (VSMC) to deterministic-policy inference. Two adjustments to the vanilla SMC sweep are required: (1) Deterministic policy consistency — For each particle, the action for a state is sampled from the proposal only on the first visit to that state and is reused on all revisits (i.e., the particle memoizes π(s)); and (2) Coupled transition randomness — "transition randomness is shared across particles within a sweep. Specifically, if two particles visit the same state s and take the same action a on the same visit count k, they are forced to transition to the same successor state s′." The optimization objective is L = log Ẑ + Σ t=1 H (log Ẑ t · Σ i=1 N log q(a t,i s t,i)), where log Ẑ t denotes the contribution of steps t through H to log Ẑ, and the overline denotes a stop-gradient operation. Theorem 1 establishes that the gradient of the surrogate objective in Eq. (6) is an unbiased estimator of ∇θ J(θ) where J(θ) = E T̂∼T, a∼qθ[log Ẑ(T̂, a)].

For policy selection, the paper states: "Action selection is performed by sampling from the posterior predictive distribution, corresponding to recurrent Thompson sampling: at each decision point, a policy is drawn from the posterior and the action prescribed by that policy is executed. The paper emphasizes that posterior predictive control yields an optimal stochastic policy under preference uncertainty, in contrast to MAP policy selection which collapses the posterior to a single deterministic behavior."

The paper distinguishes its approach from related work: "the policy itself is the latent random variable, and its expected return defines an unnormalized log density. This yields a posterior over policies directly, without adding observation channels or trajectory-level optimality variables. Compared to entropy-regularized RL, stochasticity here reflects uncertainty over deterministic policies and is realized by posterior predictive (Thompson-style) sampling, rather than being embedded as entropy inside a single learned stochastic policy."

Experiments were conducted on grid worlds, Blackjack, Triangle Tireworld, and Academic Advising, comparing VSMC to discrete Soft Actor-Critic (SAC). In grid worlds, "VSMC and SAC trajectory return distributions are close but different, and the policies differ in particular along the grid edges — SAC, optimizing an entropy-regularized stochastic policy, uses actions directed toward the grid boundaries to increase the entropy. VSMC penalizes such actions strongly because a deterministic policy directing the agent into a grid boundary can escape the current cell only due to environment stochasticity. In Blackjack, VSMC exhibits a higher expected reward than SAC with α = 1, and it takes α = 0.1 for SAC to approximately match VSMC. In Triangle Tireworld, With the original rewards, Triangle Tireworld induces a large return gap between 'fast but risky' and 'safe but slow' behaviors. Under our Bayesian formulation this makes the posterior highly peaked, yielding low mean return and high variance. Scaling rewards down by a factor of five reduces this separation, producing a less concentrated posterior; under this setting VSMC exhibits performance comparable to SAC. In Academic Advising, VSMC policy return distributions have heavier tails as manifested by 0.05 and 0.95 quantiles and their conditional tail means."

The discussion section notes: A policy posterior separates solution uncertainty from environment randomness. Policies with comparable expected return coexist in the posterior, while substantially worse policies are exponentially downweighted. Three sources of uncertainty are disentangled: aleatoric transition randomness, sampled forward and appearing as noise in the Monte Carlo estimate of policy log-probability; epistemic uncertainty over optimal behavior, represented by posterior dispersion; and execution-time stochasticity, obtained by marginalizing over deterministic policies. The paper acknowledges a limitation: "unlike classical MDP planning, which is invariant to affine reward scaling, the posterior depends on return magnitudes, so the method works best when reward scale meaningfully encodes the strength of preferences/regrets rather than merely ranking policies."

Improvements for AI systems

Based on the paper, here are specific improvements that can be made to AI systems, along with what the improved system can do:

Improvement: Replace single-policy optimization (e.g., standard RL) with posterior inference over deterministic policies, where each policy's unnormalized density is monotone in its expected return (Eq. 3). Use VSMC with the two key adaptations: (a) deterministic policy consistency via memoization on state revisits, and (b) coupled transition randomness across particles via shared cached transitions.

What the improved system can do:

  • Quantify solution uncertainty: Instead of returning one policy, the system outputs a posterior distribution over policies. This allows the agent to report I am 70% confident that going left is optimal, 30% confident that going right is optimal rather than a single action.

  • Act via recurrent Thompson sampling: At each decision point, sample a policy from the posterior and execute its action. This yields a stochastic controller that randomizes only when multiple behaviors are genuinely plausible, avoiding unnecessary randomness when the posterior is concentrated.

  • Distinguish aleatoric vs. epistemic uncertainty: The system can separate environment noise (aleatoric) from uncertainty about which policy is best (epistemic), enabling better risk assessment and more informed exploration.

Improvement: Implement the VSMC sweep with the exact algorithmic details from Algorithm 1: (i) first-visit action sampling with memoization, (ii) shared transition cache MT keyed by (s, a, k) to couple randomness across particles, (iii) adaptive resampling with ESS threshold 0.5, and (iv) the temporally stratified score-function objective in Eq. (6) with stop-gradient baselines.

Improvement: Expose the reward scale as a tunable hyperparameter that controls posterior concentration. When rewards are large, the posterior concentrates on near-deterministic optimal behavior; when rewards are small, the posterior remains diffuse, reflecting genuine preference uncertainty.

Improvement: Replace entropy-regularized stochastic policies (e.g., SAC with α=1) with posterior predictive sampling over deterministic policies. This avoids the need to tune an entropy coefficient and prevents the pathological behavior where SAC increases entropy by directing agents into grid boundaries (Figure 3b).

Improvement: Apply the VSMC framework with neural network parameterization of the proposal qθ(as) over factored state representations, using two hidden layers of width 64, 10 particles, and 50,000 iterations with cosine learning rate decay.

Improvement: Use the coupled-transition mechanism to ensure that particle weights reflect policy differences rather than independent noise realizations, which is critical in domains with irreversible failures (e.g., Triangle Tireworld where getting a flat without a spare is terminal).

Improvement: Implement the system to access the MDP only through the simulator interface Step(s, a) and Reward(s, a, s'), as specified in Eq. (1), with no need for explicit transition probability tables.


Summary of the most impactful improvement: The core contribution is replacing single-policy optimization with posterior inference over deterministic policies using VSMC with coupled randomness. This gives AI systems the ability to (1) quantify uncertainty over optimal behavior, (2) act via Thompson sampling for calibrated stochasticity, and (3) avoid the pitfalls of entropy regularization—all while maintaining the classical expected-return optimality criterion.

Abstract

We cast episodic Markov decision process (MDP) planning as Bayesian inference over policies. A policy is treated as the latent variable and is assigned an unnormalized probability of optimality that is monotone in its expected return, yielding a posterior distribution whose modes coincide with return-maximizing solutions while posterior dispersion represents uncertainty over optimal behavior. To approximate this posterior in discrete domains, we adapt variational sequential Monte Carlo (VSMC) to inference over deterministic policies under stochastic dynamics, introducing a sweep that enforces policy consistency across revisited states and couples transition randomness across particles to avoid confounding from simulator noise. Acting is performed by posterior predictive sampling, which induces a stochastic control policy through a Thompson-sampling interpretation rather than entropy regularization. Across grid worlds, Blackjack, Triangle Tireworld, and Academic Advising, we analyze the structure of inferred policy distributions and compare the resulting behavior to discrete Soft Actor-Critic, highlighting qualitative and statistical differences that arise from policy-level uncertainty.

Sources

Related papers