Safe Learning of Adaptive Control Policies for Remote Patient Monitoring
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Safe Learning of Adaptive Control Policies for Remote Patient Monitoring".
Rosa: The gist Remote Patient Monitoring (RPM) enables continuous observation of patients in their daily environments, improving both health outcomes and quality of life,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, to recap, this paper tackles the challenge of finding an optimal way to monitor patients remotely when you don't know all the underlying system details.
Dev: The core idea is to build a learning-based control framework that estimates those unknown parameters and adapts the monitoring policy in real time while balancing patient safety and monitoring costs.
Taro: They frame this as an online reinforcement learning problem where they're constantly updating their policy based on the data they collect during remote patient monitoring.
Rosa: They model the system using a Markov Decision Process with a joint state that includes both the health of the patient and whether they are under ordinary or intensive monitoring.
Dev: The thesis is that this approach allows for real-time adaptation to changing things like patient data, clinical guidelines, and institutional constraints without needing prior knowledge of those rules.
Taro: They provide theoretical guarantees on safety and convergence to the optimal policy because they explicitly prioritize patient safety during the exploration phase of their learning algorithm.
Rosa: In terms of what matters for a listener, this means we can have monitoring systems that are flexible enough to handle real-world changes without needing constant manual reprogramming.
Dev: They claim that simulation results demonstrate convergence to the optimal threshold-based policy and show that the system maintains low treatment costs while effectively reducing the risk of patients reaching critical health states.
Taro: The practical implication is that this could lead to more flexible and deployable RPM tools for clinicians because the resulting policies are designed to be transparent and interpretable.
Rosa: It's about making remote monitoring smarter by letting the system learn how to best monitor, rather than relying on a fixed setup.
Dev: And they show that while they have theoretical guarantees on safety, they also provide concrete examples of how this framework works in practice during these simulations.
Conclusion: Rosa: So, thinking about the title, "Safe Learning of Adaptive Control Policies for Remote Patient Monitoring," it really boils down to making these remote monitoring tools smarter and safer through learning.
Dev: The authors are Ramanan Tamizholi, Siddharth Chandak, Isha Thapa, Nicholas Bambos and David Scheinker. They developed an online model-based reinforcement learning algorithm specifically for RPM.
Taro: What this means in simple terms is that instead of setting fixed rules for monitoring intensity based on health levels right at the start, the system figures out those rules as it gathers real patient data over time.
Rosa: It’s about giving the monitoring system a brain that learns how to best decide when to ramp up or down observation based on what's happening with the patient.
Dev: The implication for someone just listening is that this moves us toward more flexible remote care where the monitoring strategy isn't static but can evolve alongside the patient's needs.
Taro: It also suggests that we can build in mechanisms to ensure that even while learning, the system stays within safe bounds, avoiding dangerous situations during those early stages.
Rosa: So it’s about creating a dynamic system where safety and cost management are balanced dynamically as the system learns from patient interactions.
Ramanan Tamizholi, Siddharth Chandak, Isha Thapa, Nicholas Bambos, David Scheinker
Indian Institute of Science (IISc) · Department of Electrical Engineering, Stanford University · Department of Management Science & Engineering, Stanford University · School of Medicine, Stanford University
eess.SY, cs.SY
Submitted: 2026-10-07
Updated: 2026-10-07
The gist: The gist Remote Patient Monitoring (RPM) enables continuous observation of patients in their daily environments, improving both health outcomes and quality of life, while this paper develops a
Key concepts
- Markov Decision Process (MDP)
- The system is modeled as an MDP where the state includes both the patient's health and the current monitoring setting. This mathematical framework helps define possible actions, resulting in a sequence of states, and calculating costs associated with those transitions to find the best long-term strategy.
- Optimal Monitoring Control
- This involves finding a control policy that minimizes the total expected discounted cost incurred by the patient until they reach a critical health state. The solution often takes the form of a threshold-based policy, which dictates when to switch between ordinary and intensive monitoring based on the patient's health level.
- Online Reinforcement Learning
- Because system parameters are often unknown, the control policy is updated continuously using samples gathered during remote monitoring. The algorithm operates in epochs: collecting data, updating parameter estimates, computing a new policy, and then deploying that policy to gather more data.
Terminology
Summary
The gist Remote Patient Monitoring (RPM) enables continuous observation of patients in their daily environments, improving both health outcomes and quality of life, while this paper develops a learning-based control framework that estimates system parameters and adapts monitoring policies in real time to balance patient safety and monitoring costs <ref:2610.10720#pg2>.
The Remote Patient Monitoring Model
The system is modeled as a Markov Decision Process (MDP) where the joint state is defined as the monitoring-patient joint state st:= (mt, ht) ∈ M × H =: S <ref:2610.10720#pg3>. The health state h ∈ H = 0, 1, 2, 3, H corresponds to patient health where state 0 is critical and leads to service cessation <ref:2610.10720#pg4>. Costs are interpreted from multiple perspectives; under ordinary monitoring (o), the patient incurs a constant cost Co ≥ 0 at any state (o, h) with h ∈ H > 0, and under intensive monitoring (i), the patient incurs a constant cost Ci ≥ 0 at any state (i, h) <ref:2610.10720#pg4>. Critical health states incur a significant cost Cc <ref:2610.10720#pg4>. The transition probabilities are defined based on the monitoring action and health state; for instance, under ordinary monitoring with no switching (o, h) a=o, the transition is to (o, min(h + 1, H)) with probability λo and to (o, h - 1) with probability µo = 1 − λo <ref:2610.10720#pg4>. Assumption 1 states that the transition probabilities satisfy λi ≥ λo and the costs satisfy 0 ≤ Co ≤ Ci ≤ Cc <ref:2610.10720#pg4>.
Optimal Monitoring Control
The problem is studied under the discounted cost setting of dynamic programming (DP) methodology, where costs are discounted by a factor of γ t with 0. The goal is to find an optimal control π∗ which minimizes the expected discounted cost E[h T X−1 t=0 γ t c(st, at) + γ T Cc] s0 = s <ref:2610.10720#pg4>. This leads to a dynamic programming equation where V∗(s) is the total expected (discounted) cost the patient will incur until reaching the critical state <ref:2610.10720#pg5>. The optimal policy π∗(s) is found by minimizing c(s, a) + γ X s'∈S P s's, a V∗(s') <ref:2610.10720#pg5>. A stationary monitoring policy πt,h¯ is defined by a threshold h¯ such that the patient stays in ordinary monitoring only when their health is better than h¯ <ref:2610.10720#pg4>. Theorem 1 states that under sufficient conditions, the optimal policy π∗ is a threshold-based policy πt,h¯ for some 0 ≤ h¯ ≤ H <ref:2610.10720#pg5>.
Learning the Unknown RPM Model
Since system parameters are often unknown, the problem is formulated as an online reinforcement learning (RL) problem where the control policy is continually updated based on samples observed thus far <ref:2610.10720#pg4>. The algorithm operates in epochs, where at the start of each epoch, system parameters are updated based on collected samples, a policy is computed, and then this control is used in remote monitoring of patients to obtain new samples <ref:2610.10720#pg5>. Algorithm 1 introduces a simple randomized exploration scheme tailored to the RPM setting that explicitly prioritizes patient safety during exploration <ref:2610.10720#pg5>. The deployed threshold h˜k is computed based on the current parameter estimates hˆ k, with adjustments made according to a probability sequence ρk to ensure sufficient exploration <ref:2610.10720#pg5>.
Safety Considerations and Convergence
To ensure safety while learning, the algorithm is initialized by carefully choosing priors λa,p and Ca,p and the prior strength n0 <ref:2610.10720#pg5>. Theorem 3 shows that if the prior strength satisfies n0 ≥ Ω(log(1/δ)), then with high probability, the deployed threshold is never lower than the optimal threshold throughout the learning process <ref:2610.10720#pg6>. Furthermore, Theorem 4 proves that under certain assumptions, there exists an almost surely finite random epoch KG such that for all k ≥ KG, the deployed threshold by Algorithm 1 is equal to the optimal threshold h⋆ <ref:2610.10720#pg9>. Lemma 2 shows that parameter estimates converge to the true parameter vector almost surely, and there exists a finite random epoch KG such that hˆ k = h⋆ for all k ≥ KG <ref:2610.10720#pg9>.
Performance Evaluation
Simulations evaluate the performance and safety aspects of the algorithm through Markov chain simulations, observing convergence to the optimal or near-optimal policy within a few epochs in almost all runs <ref:2610.10720#pg5>. The results show that across initial thresholds, the algorithm converges to the optimal policy (h¯ = 12) after approximately the same number of epochs, indicating robustness to initialization <ref:2610.10720#pg5>. For sufficiently large prior strengths, the average time to reach the critical health state remains consistently high throughout learning, indicating that the algorithm maintains patient safety while adapting the monitoring policy <ref:2610.10720#pg5>. The paper concludes by developing an online model-based reinforcement learning approach for optimizing RPM policies when system parameters are initially unknown <ref:2610.10720#pg7>.
Extensions
An important next step is to implement the proposed algorithm on real-world patient data to assess its practical performance <ref:2610.10720#pg7>. The paper focuses on the special case where RPM parameters are independent of health states, noting that extending this algorithm to health-state-dependent parameters is practical only for small state spaces due to the curse of dimensionality <ref:2610.10720#pg7>.
REFERENCES
[1] F. A. C. d. Farias, C. M. Dagostini, Y. d. A Bicca, V. F Falavigna, and A Falavigna, “Remote patient monitoring: a systematic review,” Telemedicine and e-Health, vol 26 no 5 pp 576–583 2020 <ref:2610.10720#pg2>.
[4] I. Lee, D. Probst, D. Klonoff, and K Sode, “Continuous glucose monitoring systems-current status and future perspectives of the flagship technologies in biosensor research,” Biosensors and Bioelectronics vol 181 p 113054 2021 <ref:2610.10720#pg5>.
[6] B. W Heckman, A R Mathew, and M J Carpenter, “Treatment burden and treatment fatigue as barriers to health,” Current opinion in psychology vol 5 pp 31–36 2015 <ref:2610.10720#pg4>.
[7] S Chandak, I Thapa, N Bambos, and D Scheinker, “Tiered service architecture for remote patient monitoring,” in 2024 IEEE International Conference on E-health Networking Application & Services (HealthCom) 2024 pp 1–7 <ref:2610.10720#pg4>.
[8] ——, “Optimal control for remote patient monitoring with multidimensional health states,” in ICC 2025 - IEEE International Conference on Communications 2025 pp 3186–3192 <ref:2610.10720#pg4>.
[9] I Thapa, P.-A Laforcade, F K Bishop, J Ferstad, M Desai, D M Maahs, P Prahalad, D Zaharieva, D Scheinker, and R Johari “Regression discontinuity in time: Evaluating the impact of evolving digital health interventions” International Journal of Medical Informatics vol 204 p 106050 2025.
[14] L N Steimle and B T Denton, Markov Decision Processes for Screening and Treatment of Chronic Diseases Cham: Springer International Publishing 2017 pp 189–222 <ref:2610.10720#pg4>.
Improvements for AI systems
-
textbfReal-time Parameter Estimation and Policy Adaptation in Unknown Environments (Algorithm 1): The system can
simultaneously learn the system parameters and adapt the monitoring policy in real time
by using an online model-based reinforcement learning algorithm that updates estimates of transition probabilities and costs based on observed patient outcomes. -
textbfSafety-Guaranteed Exploration Scheme: Prior-Aware Threshold Adjustment: The algorithm employs a
simple randomized exploration scheme tailored to the tiered RPM setting,
where it adjusts the deployed threshold based on whether the computed optimal threshold is outside the range of safe limits, such thatexploration is performed by adjusting the threshold rather than arbitrary actions, naturally supports patient-safety guarantees.
-
textbfGuaranteed Convergence to Optimal Policy: Robustness Under Uncertainty: The system can be guaranteed convergence because
with probability at least 1 − δ, the deployed policy is at least as safe as the optimal policy at every epoch,
and under specific conditions,the deployed threshold is equal to the computed threshold hˆk, and both are equal to the optimal threshold h⋆.
-
textbfClinician-Actionable Policy Output: Interpretable Threshold Policies: The resulting control policies are
transparent and interpretable,
allowing clinicians to receive a policy where decisions are clearly defined by a single threshold, such as "ordinary monitoring is used for h > 3 and intensive for h ≤ 3" in the example. -
textbfAdaptive Monitoring Intensity Balancing: Optimized Cost-Benefit Decisions: The system can dynamically balance health outcomes and costs by learning parameters to ensure the policy
maintains low treatment costs, and reduces the risk of patients reaching critical health states.
Sources
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation