Recurrent Deep Reinforcement Learning for Chemotherapy Control under Partial Observability
cs.LG, cs.AI
Submitted: 2026-05-04
Updated: 2026-05-04
Comments: Accepted for publication at the VI. International Conference on Electrical, Computer and Energy Technologies (ICECET 2026)
Journal ref: 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)
DOI: 10.1109/ICECET65726.2026.11632482
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chemotherapy dose optimization can be formulated as a dynamic treatment regime, requiring sequential decisions under uncertainty that must balance tumor suppression against toxicity.
Terminology
Abstract
Chemotherapy dose optimization can be formulated as a dynamic treatment regime, requiring sequential decisions under uncertainty that must balance tumor suppression against toxicity. However, most reinforcement learning approaches assume full observability of the patient state, a condition rarely met in clinical practice. We investigate whether memory-augmented policies can improve chemotherapy control under partial observability. To this end, we employ a recurrent TD3-based approach with separate LSTM actor-critic networks and evaluate it on the AhnChemoEnv benchmark from DTR-Bench, considering both off-policy and on-policy recurrent architectures against feed-forward TD3 and Soft Actor-Critic. Pharmacokinetic and pharmacodynamic variability are held fixed to isolate hidden-state uncertainty and observation noise and to avoid confounding effects from inter-patient variability. Across ten random seeds, recurrence yields modest benefit under full observability but substantially stronger and more stable performance under partial observability, with more consistent tumor suppression and improved normal-cell preservation. These findings indicate that memory-based policies are particularly beneficial when clinically relevant state information is incomplete or noisy.
Sources
- Continuous control with deep reinforcement learning
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs
- Addressing Function Approximation Error in Actor-Critic Methods
- Proximal Policy Optimization Algorithms
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks