Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
math.OC, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The growing demand for real-time, data-driven decision-making in complex and dynamic systems is placing increasing pressure on traditional Operational Research (OR) methodologies.
Terminology
Abstract
The growing demand for real-time, data-driven decision-making in complex and dynamic systems is placing increasing pressure on traditional Operational Research (OR) methodologies. Reinforcement learning (RL) has emerged as a complementary approach, offering strong learning and computational capabilities for sequential decision-making in dynamic and uncertain environments. Recent research shows an increasing interest in integrating RL with OR to address dynamic decision-making problems, enhance heuristic and exact methods for combinatorial optimization, and support the development of digital replicas of operational systems. The overarching goal across these efforts is to leverage the learning capabilities of RL to strengthen traditional OR algorithms, improving solution quality, computational efficiency, and robustness. Given the diversity of integration approaches and application settings, there is a clear need for a systematic and technically detailed review of how RL empowers OR methods. To address this gap, this paper presents a structured review of three key roles that RL plays in empowering OR: (i) solving sequential decision-making problems in dynamic environments, (ii) serving as an end-to-end solution method or as a component integrated within heuristic and exact OR methods for combinatorial optimization problems, and (iii) facilitating extended reality analysis through integration with digital twin systems. We critically synthesize recent advances across these roles, highlighting their advantages, implementation requirements, limitations, and challenges. Finally, based on these insights, we outline a roadmap for future research to further advance the methodological and practical integration of RL and OR.
Sources
- Neural Combinatorial Optimization with Reinforcement Learning
- Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Deep Reinforcement Learning
- Continuous control with deep reinforcement learning
- Asynchronous Methods for Deep Reinforcement Learning
- Playing Atari with Deep Reinforcement Learning
- Reinforcement Learning: An Overview
- Is Value Learning Really the Main Bottleneck in Offline RL?
- Trust Region Policy Optimization
- Proximal Policy Optimization Algorithms
- VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification