Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers
math.OC, cs.AI, cs.LG, cs.SY, eess.SY, stat.ML
Submitted: 2026-04-13
Updated: 2026-09-15
Journal ref: I. Esra Buyuktahtakin. Deep learning for sequential decision making under uncertainty: Foundations, frameworks, and frontiers. Forthcoming in INFORMS TutORials in Operations Research, 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Artificial intelligence (AI) is moving increasingly beyond prediction to support decisions in complex, uncertain, and dynamic environments.
Terminology
Abstract
Artificial intelligence (AI) is moving increasingly beyond prediction to support decisions in complex, uncertain, and dynamic environments. This shift creates a natural intersection with operations research and management science (OR/MS), which has long provided methodological foundations for sequential decision making under uncertainty. At the same time, deep learning advances, including feedforward neural networks, recurrent architectures, transformers, large language models (LLMs), and deep reinforcement learning, have expanded data-driven modeling for large-scale decisions. This tutorial presents an OR/MS-centered perspective on deep learning for sequential decision making under uncertainty, bridging neural architectures and OR/MS approaches to decision making. Its premise: deep learning complements optimization rather than replacing it. Deep learning brings adaptability and scalable approximation, whereas OR/MS provides the mathematical rigor to represent constraints, recourse, uncertainty, and decision quality. The tutorial reviews key decision making foundations, connects them to the major neural architectures in modern AI, and organizes the field around three central themes: predict-then-optimize and decision-aware learning, learning-based decision generation under constraints for continuous and discrete problems with temporal coupling, and deep reinforcement learning for sequential and combinatorial decision making. Impact spans supply chains, service systems, healthcare and epidemic response, agriculture, energy, environmental sustainability, and autonomous operations. This tutorial frames these developments as part of a shift from predictive AI toward decision-capable AI, highlighting OR/MS's role in shaping the next generation of integrated learning--optimization systems.
Sources
- OptNet: Differentiable Optimization as a Layer in Neural Networks
- OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Language Models are Few-Shot Learners
- Toward TransfORmers: Revolutionizing the Solution of Mixed Integer Programs with Transformers
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Exact Combinatorial Optimization with Graph Convolutional Neural Networks
- Inductive Representation Learning on Large Graphs
- Deep Recurrent Q-Learning for Partially Observable MDPs
- A Multi-echelon Demand-driven Supply Chain Model for Proactive Optimal Control of Epidemics: Insights from a COVID-19 Study
- Learning Combinatorial Optimization Algorithms over Graphs
- Semi-Supervised Classification with Graph Convolutional Networks
- Attention, Learn to Solve Routing Problems!
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- A Unified Approach to Interpreting Model Predictions
- Reinforcement Learning for Solving the Vehicle Routing Problem
- Neur2SP: Neural Two-Stage Stochastic Programming
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Certifying Some Distributional Robustness with Principled Adversarial Training
- Sequence to Sequence Learning with Neural Networks
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification