Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control

arXiv:2608.10777 · cs.LG, math.OC, stat.ML · Submitted 2026-08-11 · Read on arXiv

Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu

Westlake University · School of Engineering

cs.LG, math.OC, stat.ML

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Project Page: https://github.com/bangyan101/PIVM/

Code: https://github.com/bangyan101/PIVM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper addresses the computational challenges in solving Linear Quadratic Stochastic Optimal Control (LQ-SOC) problems.

Terminology

Summary

The paper addresses the computational challenges in solving Linear Quadratic Stochastic Optimal Control (LQ-SOC) problems. The authors state: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. The formulation is defined as:

u in U E P u [integral 0 1 (1 over 2u(X t,t) squared + f(X t,t)) dt + g(X 1)]

subject to the controlled SDE: dX t = (b(X t,t) + sigma(t)u(X t,t))dt + sigma(t)dB t.

The authors identify a critical bottleneck in existing approaches: "Current state-of-the-art approaches largely follow the policy-based paradigm, known as the Iterative Diffusion Optimization. These methods optimize neural control policies by minimizing the divergence between parameterized path measures and optimal path measures. However, a critical bottleneck restricts their scalability: they rely heavily on on-policy full-trajectory simulation for gradient estimation. This dependence on long-horizon stochastic simulation leads to excessive computational costs and high-variance gradient estimates in high dimension."

The central theoretical insight is a recursive formulation of the path integral representation of the optimal value function. The paper establishes Proposition 3.1: For any 0 < t < s, the following expectation equation holds:

(-V(X t,t)) = E P 0 [(-V(X s,s)) (-integral t s f(X r,r)dr) F t]

with terminal condition V(x,1) = g(x). This is derived by decomposing the path integral in Equation (5) into the intervals [t, s] and [s, 1], and taking the conditional expectation over the trajectory from s to the terminal time.

Theorem 3.3 establishes that the iterative update rule converges to the optimal value function under assumptions including: The running cost is bounded from below by a positive constant; The optimal value function is bounded from below; and The underlying SDE satisfies standard Lipschitz and growth conditions, ensuring the transition semigroup possesses the Feller property. The proof uses the Banach Fixed Point Theorem by showing the update operator is a contraction mapping with contraction factor gamma = e-f inf epsilon < 1.

Proposition 3.4 provides a variance decomposition showing the explicit variance reduction achieved by the recursive approach:

V(Z t F t) = V (E[Z s F s] (-integral t s f(X r,r)dr) F t) + E [V(Z s F s) (-2 integral t s f(X r,r)dr) F t]

The authors explain: "The term on the left-hand side (LHS) represents the variance of the naive Monte Carlo estimator. The first term on the right-hand side (RHS) corresponds to the variance of our recursive estimator... Consequently, the second term on the RHS quantifies the explicit variance reduction achieved by our method. This reduction is particularly substantial when the current time t is far from the terminal time."

The paper proposes the Path Integral Value Matching (PI-VM) algorithm, which employs Temporal Difference (TD) learning to approximate the recursive value dynamics, effectively replacing high-variance long-horizon integration with stable short-term bootstrapping.

The Path Integral TD Loss (Definition 4.1) is defined as:

(theta) = (V theta(x,t) - theta(x,t,s)) squared

where the target value is estimated via Monte Carlo approximation:

theta(x,t,s) = - (1 over N sum j=1 N (- (Xt,s) - (X s(j))))

The paper extends the method to off-policy training: "We integrate the Experience Replay buffer with the Girsanov theorem to support off-policy training. This combination allows for mathematically grounded trajectory reweighting, enabling flexible transitions between on-policy exploitation and off-policy exploration."

The Off-Policy PI-VM Loss (Definition 4.2) replaces the uncontrolled target with a controlled one using the Radon-Nikodym derivative:

dP 0[t,s] over dP u[t,s] = (-integral t s u(X r,r)dB r - 1 over 2 integral t s u(X r,r) squared dr)

Theorem 4.3 provides a variance bound for the off-policy estimator: "Suppose that the control u satisfies x,ru*(x,r) - u(x,r) squared at most kappa. Then, the conditional variance of M(X[t,s]) satisfies the following upper bound: V P u(M(X[t,s]) F t) at most (-2V(X t,t))[(kappa(s-t)) - 1]. This shows the variance bound diminishes as a function of the control error κ. As the current control u converges toward u*, the upper bound becomes increasingly tighter, ensuring that the estimation variance asymptotically vanishes."

The paper evaluates on three unimodal tasks: Linear Ornstein-Uhlenbeck, Quadratic Ornstein-Uhlenbeck (easy), and Quadratic Ornstein-Uhlenbeck (hard). Results show: "While all methods perform comparably on the simple Linear task, our PI-VM method significantly outperforms baselines on the Quadratic tasks. Notably, in the 'Hard' setting, state-of-the-art methods like SOCM and SOCM-A fail to converge. In contrast, our value-based approach successfully approximates the global landscape, achieving the lowest error. Furthermore, PI-VM demonstrates superior efficiency, running 10–20× faster than baselines."

For the Gaussian Mixture Model (GMM) task in 20 dimensions: "our method consistently outperforms the baseline across all configurations. Notably, our approach demonstrates superior robustness in scenarios with Small Variance (i.e., highly peaked distributions). In these challenging settings, the baseline method suffers from catastrophic failure."

On the Quadratic OU Easy task with dimensions increasing from 10 to 200: "The results reveal severe limitations in baseline methods: SOCM quickly encounters memory bottlenecks (OOM) at d = 80 due to the prohibitive cost of storing full trajectories, while AM suffers from optimization instability, resulting in error explosion and eventual divergence at d = 200. In stark contrast, ours PI-VM effectively breaks the curse of dimensionality, maintaining robust convergence and real-time inference speeds even at d = 200."

On the 50-dimensional Many-well potential landscape: "our learned Value Function accurately reconstructs the underlying non-convex geometry, providing a correct guidance for control... the Adjoint Matching baseline suffers from high variance, with a significant number of samples failing to converge and scattering across high-energy barriers. In contrast, our method generates high-fidelity samples that are strictly confined to the stable equilibria."

The default configuration uses N=8 Monte Carlo samples and M=8 forward steps: The empirical results suggest that the configuration (N = 8, M = 8) offers the optimal trade-off between accuracy and inference speed. Additional ablations show deeper networks achieve lower errors, and larger batch sizes with moderate sample sizes perform best.

The paper emphasizes fundamental differences from existing approaches:

  • PI-VM vs. classical RL: Directly taking the discrete-to-continuous limit of these equations does not yield our log-sum-exp formulation. Instead, our formulation arises directly from path integral control.

  • PI-VM vs. risk-sensitive RL: Risk-sensitive RL applies it to the value function, whereas in our method, the log-sum-exp formulation applies to the optimal value function.

The authors acknowledge: As a value-based approach, PI-VM still requires computationally expensive automatic differentiation to obtain the control signal, which can limit runtime efficiency in practice.

The paper concludes: "We presented PI-VM, a novel value-based algorithm for high-dimensional Stochastic Optimal Control. By revisiting Path Integral Control and deriving a temporal recursive formulation, we effectively eliminated the need for computationally expensive and high variance full-trajectory simulations. Our method leverages TD learning, experience replay buffer, and Girsanov theorem to achieve stable and efficient off-policy training. Experimental results confirm that PI-VM significantly outperforms existing policy-based baselines in terms of training speed, numerical stability, and scalability to high-dimensional state spaces."

Improvements for AI systems

Improvements to AI Systems:

  1. Recursive Value-Based Control for Long-Horizon Tasks

Replace policy-gradient or full-trajectory simulation methods with a temporal-difference (TD) framework that bootstraps value estimates over short intervals. This reduces variance and computational cost, enabling stable training for high-dimensional control (e.g., robotics, autonomous navigation) without requiring full episode rollouts.

  1. Off-Policy Learning with Girsanov Reweighting

Integrate experience replay with mathematically grounded trajectory reweighting (via Radon–Nikodym derivatives) to decouple exploration from exploitation. This allows AI agents to reuse past data efficiently, improving sample efficiency and enabling safe exploration in real-world systems (e.g., drone flight, chemical process control).

  1. Variance-Controlled Estimators for Stochastic Optimization

Adopt the recursive variance decomposition (Proposition 3.4) to design estimators with provably lower variance than naive Monte Carlo. This improves convergence speed and stability for any stochastic optimization or reinforcement learning task, particularly in high-dimensional spaces where naive sampling fails.

  1. Scalable Non-Convex Landscape Approximation

Use the value-matching loss to learn global value functions that capture non-convex cost landscapes (e.g., multi-modal distributions). This enables AI systems to perform robust mode coverage and avoid local optima, useful for generative modeling, molecular conformation sampling, and multi-objective planning.

  1. Real-Time Inference with Memory Efficiency

Replace full-trajectory storage with short-horizon bootstrapping, reducing memory footprint (e.g., from O(d×T) to O(d×M) with M≪T). This allows AI systems to run on edge devices or in real-time applications (e.g., autonomous vehicles, financial trading) with limited computational resources.

  1. Theoretical Guarantees for Iterative Value Updates

Leverage the contraction mapping proof (Theorem 3.3) to ensure monotonic convergence of value function approximations. This provides reliability for safety-critical AI systems (e.g., medical treatment planning, power grid control) where non-convergent algorithms are unacceptable.

  1. Adaptive Control Error Bounding

Use the off-policy variance bound (Theorem 4.3) to dynamically adjust exploration noise or control policies, ensuring estimation variance remains bounded as the policy improves. This enables self-tuning AI agents that balance exploration and exploitation without manual hyperparameter tuning.

What the Improved AI System Can Do:

  • Solve high-dimensional stochastic optimal control problems (e.g., 200+ state dimensions) in real time, with 10–20× faster training than current policy-based methods.

  • Learn from off-policy data (e.g., previously collected trajectories) without performance collapse, enabling continuous learning in changing environments.

  • Generate high-fidelity samples from complex, multi-modal distributions (e.g., for generative AI or Bayesian inference) with guaranteed convergence and low variance.

  • Operate on resource-constrained hardware (e.g., embedded systems) by avoiding full-trajectory memory bottlenecks.

  • Provide provably stable value function learning for safety-critical applications, with explicit control over estimation error during training.

Abstract

Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. However, current state-of-the-art policy-based methods suffer from prohibitive computational costs and instability due to their heavy reliance on full-trajectory simulation. To overcome these limitations, we propose a paradigm shift toward a value-based approach by revisiting Path Integral Control (PIC). Although standard PIC suffers from the same high-variance bottleneck as policy-based methods, we discover that by truncating and marginalizing the original path integral formulation, we can derive a temporal recursive form of the value function. Building upon this theoretical foundation, we propose the Path Integral Value Matching (PI-VM) algorithm. Specifically, we employ temporal-difference learning to approximate the recursive value dynamics, and further integrate the Girsanov theorem with experience replay to enable off-policy training. We benchmark PI-VM against SOTA policy-based methods across various SOC benchmarks and sampling tasks. Empirical results demonstrate that PI-VM matches SOTA precision with an order-of-magnitude efficiency gain in low-dimensional settings, while effectively mitigating mode collapse in high-dimensional scenarios. Consequently, PI-VM offers a scalable solution for solving complex SOC problems.

Sources

Related papers