Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

arXiv:2609.26068 · cs.NI, cs.LG · Submitted 2026-08-02 · Read on arXiv

cs.NI, cs.LG

Submitted: 2026-08-02

Updated: 2026-08-02

License: http://creativecommons.org/licenses/by/4.0/

The gist: Learning-based control is increasingly central to automating network operations.

Terminology

Abstract

Learning-based control is increasingly central to automating network operations. A learned policy, however, must satisfy cross-layer constraints on interference, power-rate coupling, flow conservation, service chains, capacity, latency, and reliability. Existing methods typically account for only a subset of this geometry and only indirectly, e.g., through reward penalties, Lagrange multipliers, or post-hoc repairs. This paper proposes Network Feasibility Geometry Reinforcement Learning (NFG-RL), which models coupled constraints via transport theory and the residual inclusion N(x,a) in N, defining the executed policy as the pushforward of a proto-policy through a feasibility-transport map. NFG-RL compiles heterogeneous constraints into typed residual blocks and transports proto-actions through a differentiable variational operator, letting active constraints shape execution, exploration, and actor gradients. Our analysis shows that exact transport yields almost-sure feasible execution, while active constraints contract exploration onto the feasible tangent space. It further establishes a nonnegative first-order gain from critic-tilted transport over plain projection and recovers backpressure scheduling as the gradient of a lifted drift residual. In two public-trace-conditioned wireless-edge surrogate environments, NFG-RL improves feasible utility by 37.5--41.5% over the strongest non-NFG method in each environment, reduces raw-action violation by 48.5--60.8%, and lowers P99 delay by 57.0--75.5%, outperforming a range of optimization and learning baselines.

Sources

Related papers