Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints

arXiv:2603.07237 · eess.SY, cs.SY · Submitted 2026-03-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Reinforcement Learning for Vehicle-to-Grid Voltage Regulation".

Dev: This paper presents a Vehicle-to-Grid (V2G) coordination framework using reinforcement learning (RL).

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, let's start by looking at the title and the authors of "Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints." I want to explain what that means in plain terms for our audience.

Dev: I think the title immediately signals that this research isn't just about making an EV charge smarter; it’s about a complex control strategy that coordinates multiple charging points across a whole system, which is a significant step up from single-point management.

Taro: I see the phrase "Battery-Aware Constraints" appearing right away, and that tells me the authors are serious about incorporating the real limitations of the batteries into their design, which is usually where these things fall short in theoretical papers.

Rosa: That’s right; it means they aren't just throwing a general RL agent at the problem; they are building a system that understands how much energy a battery can safely put into or take from the grid based on its current health and charge level.

Dev: And the coordination aspect, single-hub to multi-hub, suggests they are addressing the challenge of scaling control efforts up to handle larger distribution feeders where one hub simply isn't enough for global stability.

Taro: It seems like they are tackling the problem from two angles: optimizing the local action at each hub and then making sure those actions work together to satisfy a larger system-wide goal, which is a complex autonomy challenge.

Rosa: Exactly, Taro; it’s about achieving that balance between local optimization and global system stability while respecting all the physical limitations imposed by the fleet resources involved.

Dev: And from an engineering view, having them define a specific state space S consisting of bus voltage magnitudes V pu i for all monitored buses in per unit (p.u.) gives us a very concrete idea of the input the RL agent is actually processing.

Taro: That large state space makes me wonder about computational feasibility; how does their chosen SAC algorithm manage that complexity without becoming too slow for real-time decision-making?

Rosa: I'm curious if the entropy term in the SAC formulation helps manage that exploration effectively, ensuring the agent finds a stable and practical policy rather than just wandering around trying random actions.

The paper's summary: Rosa: Now that we’ve talked about the title, let’s get into the main summary of "Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints." Essentially, what is the core contribution they are making?

Dev: The core contribution is a reinforcement learning coordination framework using a soft actor-critic algorithm designed to regulate voltage in distribution networks through both single and multi-hub charging systems while strictly adhering to fleet constraints.

Taro: So, they’ve developed an intelligent control strategy that uses RL to manage the power setpoints at these V2G hubs, with the key challenge being ensuring that the actions taken by all hubs collectively maintain voltage within safe limits.

Rosa: They also highlight a two-phase training approach designed to integrate stability-focused learning with battery-aware deployment, which is meant to ensure the final result is practically feasible for real operation.

Dev: In simulation studies on the IEEE thirty-four-bus system, they validate this framework against a standard Volt-Var/Volt-Watt droop controller, showing that the RL agent achieves performance comparable to the baseline control strategy in nominal scenarios.

Taro: The real test, according to their methodology, is how it performs under aggressive overloading conditions where it needs to provide robust voltage recovery while simultaneously prioritizing fleet availability and state-of-charge preservation.

Rosa: It sounds like the summary boils down to a system that learns an optimal control policy using RL, which is then rigorously tested against established local control strategies under both normal and extreme conditions.

Dev: The methodology involves a hierarchical allocation mechanism where a fleet-aware power mapping module translates those hub-level signals into physically realizable battery actions by incorporating factors like power conversion, inverter efficiency, and SOC/SOH-dependent bounds.

Taro: That mapping layer is critical because it’s the bridge between the abstract RL decision and the physical reality of the battery system; if that module fails or miscalculates, everything else in this framework collapses.

Rosa: So, to put it simply, they’ve built a comprehensive system that learns how to manage voltage across multiple V2G hubs intelligently while ensuring the learned actions are physically possible given the batteries and fleet availability.

The paper's improvements: Rosa: Moving on to the specific improvements suggested by the authors of "Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints," they propose a two-phase training approach as a major advancement.

Dev: They suggest this two-phase workflow—Phase one trains the agent in an idealized environment with fixed hub power limits and no explicit fleet constraints, using time-varying load conditions through load multipliers lambda in

lambda min, lambda max: .

Taro: That separation is smart because it lets the agent learn the fundamental stability policies first without getting bogged down by the complexity of real-world fleet logistics during that initial learning phase.

Rosa: Exactly, Taro; Phase one establishes a foundation in stability using fixed limits, and then Phase two deploys that policy while enforcing dynamic constraints like SOC and SOH evolution according to SOC e(t + t) = SOC e(t) + I bat e t / C eSOHe.

Dev: The second phase then incorporates the detailed fleet model, where a hub-level scaling ratio rho(h) adjusts the agent's outputs based on real-time fleet availability. This allows for dynamic adaptation to changing conditions during operation.

Taro: I think incorporating time-varying EV availability schedules directly into the RL state space or reward function would really push this toward a more practical system, allowing it to learn optimal charging/discharging schedules that respect vehicle travel logistics.

Rosa: That would be a huge leap; linking the learning process directly to scheduling constraints makes the resulting control strategy much more immediately applicable in a commercial setting, moving it from simulation to practical deployment.

Dev: The paper also notes that for multi-hub coordination, spatial coordination becomes essential; the improved system is designed to optimize voltage regulation across multiple geographically distributed hubs simultaneously when local hub control is insufficient.

Conclusion: Rosa: So, wrapping up our discussion on "Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints," what are the main implications we should take away from this work?

Dev: The main takeaway is that coordinated RL can provide meaningful feeder-wide support, especially when dealing with multi-hub setups under stress, though the paper does point out that a local droop baseline can still outperform it under aggressive stress.

Taro: I think the implication for autonomy research is that as grid infrastructure gets more distributed, we need intelligent, coordinated control systems like this to manage the complexity and maintain stability when things go wrong.

Rosa: I’m particularly interested in the practical application here; if these constraints are handled correctly, this framework suggests a path toward more reliable voltage management for distribution networks by balancing learned coordination with physical battery realities.

Dev: From an engineering standpoint, the two-phase training approach is a valuable lesson in how to build robust AI systems that transition successfully from theory into deployment by systematically introducing complexity.

Taro: I just think the future work needs to focus on making those constraints even more dynamic and incorporating real-time operational data to ensure this framework remains effective in unpredictable grid conditions.

Rosa: It sounds like this paper lays a solid groundwork for using RL not just for optimization, but for resilient, constraint-aware control in energy systems. That's really encouraging stuff.

The University of Texas at Dallas

eess.SY, cs.SY

Submitted: 2026-03-07

Updated: 2026-03-07

Journal ref: 2026 IEEE Power & Energy Society General Meeting (PESGM)

DOI: 10.1109/PESGM58988.2026.11693776

Code: https://github.com/dss-extensions/OpenDSSDirect.py

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: This paper presents a Vehicle-to-Grid (V2G) coordination framework using reinforcement learning (RL).

Key concepts

Vehicle-to-Grid (V2G) Coordination
This refers to a framework where multiple electric vehicles coordinate their charging or discharging actions across a distribution network. The research focuses on coordinating these actions, moving beyond simple single-point management to handle complex multi-hub systems for better grid stability.
Battery-Aware Constraints
These are the real limitations of batteries that the system must respect. This includes factors like how much energy a battery can safely put into or take from the grid based on its current charge level and health, ensuring physical feasibility in the control strategy.
Two-Phase Training Approach
This is a method for training the reinforcement learning agent. Phase one trains the agent in an idealized environment with fixed limits to learn basic stability policies first. Phase two then incorporates dynamic constraints like battery state-of-charge and real fleet availability for practical deployment.
Hierarchical Allocation Mechanism
This module is critical in translating the abstract decisions made by the RL agent into physical actions for the batteries. It incorporates factors such as inverter efficiency and bounds related to the battery's current state of charge (SOC) or state of health (SOH).

Terminology

Summary

This paper presents a Vehicle-to-Grid (V2G) coordination framework using reinforcement learning (RL). An intelligent control strategy based on the soft actor-critic algorithm is developed for voltage regulation through single and multi-hub charging systems while respecting realistic fleet constraints. A two-phase training approach integrates stability-focused learning with battery-aware deployment to ensure practical feasibility.

Simulation studies on the IEEE 34-bus system validate the framework against a standard Volt-Var/Volt-Watt droop controller. Results indicate that the RL agent achieves performance comparable to the baseline control strategy in nominal scenarios. Under aggressive overloading, it provides robust voltage recovery (within 10% of the baseline) while prioritizing fleet availability and state-of-charge preservation, demonstrating the viability of constraint-aware learning for critical grid services.

The methodology involves a hierarchical allocation mechanism where a fleet-aware power mapping module translates hub-level power signals into battery-level actions by incorporating power conversion, inverter efficiency, and SOC/SOH-dependent bounds. The RL framework is formulated as a Markov Decision Process (MDP) defined by (S, A, P, r), where the state space S consists of bus voltage magnitudes Vpu i for all monitored buses in per unit (p.u.), and the action space A contains continuous power scaling factors for hub injections, normalized control factors bounded in [-1, 1]. The reward function guides the RL model to regulate voltage in the network through the V2G control: rt = Rvb − Rvp, where Rvb is a reward based on whether all voltages are within range [0.95 p.u., 1.05 p.u.] and Rvp penalizes deviations outside this range using terms like (Vmin − Vi) × 100 if Vi Vmax, with Vmin = 0.95 p.u. and Vmax = 1.05 p.u. The agent employs the Soft Actor-Critic (SAC) algorithm to maximize the expected cumulative reward augmented with an entropy term to promote exploration: J(π) = E(st,at)∼D [Qθ(st, at) − α log πϕ(atst)], where Qθ is the critic network and πϕ is the stochastic actor network.

The training and deployment workflow employs a two-phase approach. Phase 1 trains the agent in an idealized environment with fixed hub power limits and no explicit fleet constraints, using time-varying load conditions through load multipliers λ ∈ [λmin, λmax] applied to base loads: P load i(λ) = λPbase i, Qload i(λ) = λQbase i. Phase 2 evaluates the trained policy with detailed fleet model enabled, where hub-level scaling ratio ρ(h) adjusts agent outputs based on real-time fleet availability while SOC and SOH states evolve dynamically according to SOCe(t + ∆t) = SOCe(t) + I bat e ∆t / CeSOHe. The case study is validated on the IEEE 34-bus radial distribution feeder, with evaluation scenarios including single-hub coordination (one V2G hub at bus 890) and multi-hub coordination (five V2G hubs at buses 890, 844, 832, 830, and 860). The performance is compared against a decentralized Droop control strategy implementing autonomous piecewise-linear Volt-Var (Q(v)) and Volt-Watt (P(v)) curves with a deadband of ±0.02 p.u. and saturation at [0.90, 1.10] p.u., benchmarked against both unconstrained and EV-constrained scenarios under mild and aggressive overloading conditions to assess controller capability versus realistic deployment limits, fleet availability, and SOC limitations as primary performance bottlenecks in single-hub settings, and the necessity of spatial coordination in multi-hub settings. The conclusion notes that while coordinated RL can provide meaningful feeder-wide support, a local droop baseline can still outperform it under aggressive stress.

Index Terms—Vehicle-to-grid, electric vehicle fleet, voltage regulation, distribution networks, reinforcement learning, soft actor-critic, multi-hub coordination.

The paper is organized as follows: Section II presents the methodology; Section III describes the case study setup and RL agent training environment; Section IV presents the simulation results and comparative analysis against industry standard Volt-Var/Volt-Watt local control strategies; and finally, Section V concludes the paper with key findings and future research directions.

Improvements for AI systems

Here are specific improvements for AI systems based on this research, focusing on the capabilities of the proposed V2G coordination framework:

  1. Improved AI System Capability: A sophisticated, constraint-aware Vehicle-to-Grid (V2G) control system capable of optimizing energy dispatch across single or multi-hub distribution networks while respecting real-world battery degradation and fleet availability constraints.

  2. Specific Improvements & Functionality:

Ease the development of V2G control algorithms by integrating a two-phase training approach: an initial phase for learning fundamental voltage regulation policies in an idealized environment (without fleet constraints) followed by a deployment phase that enforces realistic, dynamic constraints (SOC/SOH and availability). This prevents the RL agent from learning unstable or infeasible actions during real-world operation.

  1. Specific Improvements & Functionality:

Enhance the system's ability to handle heterogeneous EV fleets by incorporating a fleet-aware power mapping module. This module translates high-level hub control signals into physically realizable battery actions by dynamically calculating power limits based on individual EV State of Charge (SOC), State of Health (SOH), C-rate limitations, and inverter efficiency.

  1. Specific Improvements & Functionality:

Enable scalable grid management by implementing a coordinated multi-hub RL framework. The improved system can optimize voltage regulation across multiple geographically distributed charging hubs simultaneously, allowing the AI to perform feeder-wide voltage support when local hub control is insufficient (as validated in the aggressive loading scenario).

  1. Specific Improvements & Functionality:

Improve robustness against extreme grid stress by benchmarking RL performance against a decentralized Volt-Var/Volt-Watt droop controller. The improved system can be configured to leverage the strengths of both approaches: using RL for flexible, learned coordination under moderate conditions and relying on the robust, saturation-aware droop strategy for high-stress scenarios where rapid, localized correction is needed.

  1. Specific Improvements & Functionality:

Optimize fleet utilization by incorporating time-varying EV availability schedules (e.g., delivery windows) directly into the RL state space or reward function during deployment phase training. This allows the AI to learn optimal charging/discharging schedules that maximize grid service while respecting realistic constraints like vehicle travel logistics, leading to more practical and reliable V2G operation.

Sources

Related papers