Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors
cs.AI, eess.SP
Submitted: 2026-04-06
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reconfigurable Intelligent Surfaces (RIS) have the potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks.
Terminology
Abstract
Reconfigurable Intelligent Surfaces (RIS) have the potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks. However, the prohibitive computational overhead of Channel State Information (CSI) estimation and the dimensionality explosion inherent in centralized optimization severely hinder practical large-scale deployments. To overcome these bottlenecks, we introduce a per-element CSI-free paradigm powered by a Hierarchical Multi-Agent Reinforcement Learning (HMARL) architecture to control mechanically reconfigurable reflective surfaces. By substituting pilot-based channel estimation for each element of the device with accessible user localization data, our framework leverages spatial intelligence for macro-scale wave propagation management. The control problem is decomposed into a two-tier neural architecture: a high-level controller executes temporally extended, discrete user-to-reflector allocations, while low-level controllers autonomously optimize continuous focal points using Multi-Agent Proximal Policy Optimization (MAPPO) under a Centralized Training with Decentralized Execution (CTDE) scheme. Comprehensive deterministic ray-tracing evaluations in an indoor mmWave scenario demonstrate that this hierarchical framework achieves received signal strength indicator (RSSI) improvements of up to 7.79 dB over centralized Proximal Policy Optimization (PPO) baselines. Furthermore, the system maintains resilient beam-focusing performance under practical sub-meter localization tracking errors for up to four users and two reflector arrays. By eliminating execution-time CSI overhead while preserving high-fidelity signal redirection, this work provides a scalable and cost-effective step toward intelligent indoor wireless environments.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection