Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry
cs.NI, cs.LG
Submitted: 2026-08-02
Updated: 2026-08-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning-based control is increasingly central to automating network operations.
Terminology
Abstract
Learning-based control is increasingly central to automating network operations. A learned policy, however, must satisfy cross-layer constraints on interference, power-rate coupling, flow conservation, service chains, capacity, latency, and reliability. Existing methods typically account for only a subset of this geometry and only indirectly, e.g., through reward penalties, Lagrange multipliers, or post-hoc repairs. This paper proposes Network Feasibility Geometry Reinforcement Learning (NFG-RL), which models coupled constraints via transport theory and the residual inclusion N(x,a) in N, defining the executed policy as the pushforward of a proto-policy through a feasibility-transport map. NFG-RL compiles heterogeneous constraints into typed residual blocks and transports proto-actions through a differentiable variational operator, letting active constraints shape execution, exploration, and actor gradients. Our analysis shows that exact transport yields almost-sure feasible execution, while active constraints contract exploration onto the feasible tangent space. It further establishes a nonnegative first-order gain from critic-tilted transport over plain projection and recovers backpressure scheduling as the gradient of a lifted drift residual. In two public-trace-conditioned wireless-edge surrogate environments, NFG-RL improves feasible utility by 37.5--41.5% over the strongest non-NFG method in each environment, reduces raw-action violation by 48.5--60.8%, and lowers P99 delay by 57.0--75.5%, outperforming a range of optimization and learning baselines.
Sources
- Geometry of Drifting MDPs with Path-Integral Stability Certificates
- Metric-Gradient Projection for Stable Multi-Agent Policy Learning
- Reward Constrained Policy Optimization
- Safe Exploration in Continuous Action Spaces
- Enabling SLO-Aware 5G Multi-Access Edge Computing with SMEC
- Proximal Policy Optimization Algorithms
- Modeling Other Players with Bayesian Beliefs for Games with Incomplete Information
- Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
- Operator-Guided Invariance Learning for Continuous Reinforcement Learning
- NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic