PEARL: Structural Privacy-Utility Control in Human-Centric CPS via Personalized Early-Exit Deep Reinforcement Learning
cs.LG, cs.CR, cs.HC
Submitted: 2024-03-09
Updated: 2026-09-10
Comments: 22 pages, 13 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: In human-centric Cyber-Physical Systems (CPS), personalized Deep Reinforcement Learning (DRL) agents must share fine-grained control actions with cloud services, exposing sensitive private states to
Terminology
Abstract
In human-centric Cyber-Physical Systems (CPS), personalized Deep Reinforcement Learning (DRL) agents must share fine-grained control actions with cloud services, exposing sensitive private states to inference attacks by honest-but-curious adversaries. Static privacy models fail to address the dynamic nature of human interactions. This paper introduces PEARL (Personalized Early-exit Adaptive Reinforcement Learning), a novel framework that addresses this challenge through structural privacy control rather than data perturbation. PEARL deploys a dual-path Early-Exit Deep Q-Network (EE-DQN) at the edge, using Mutual Information (MI) between private states and observable actions to train per-branch binary labels: Utility Confidence Labels (UCL), verifying action quality, and Privacy Confidence Labels (PCL), verifying MI leakage remains below a user-defined threshold. At inference, PEARL selects the shallowest exit branch satisfying both UCL and PCL, structurally limiting shared action descriptive power without noise injection. An MI-based feedback loop tracks behavioral drift and triggers retraining when privacy-utility profiles shift, ensuring long-term robustness. Validated on a personalized smart-home HVAC system and a VR smart classroom, PEARL reduces adversarial state-inference accuracy by 25.67% on average with a controlled 10-16% utility cost, establishing a practical, dynamically enforceable privacy-utility tradeoff.
Sources
- adaPARL: Adaptive Privacy-Aware Reinforcement Learning for Sequential-Decision Making Human-in-the-Loop Systems
- FAIRO: Fairness-aware Adaptation in Sequential-Decision Making for Human-in-the-Loop Systems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks