Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings
Tran Le Vu
Energy Research Institute @ Nanyang Technological University
cs.LG, cs.AI, cs.NE, math.OC
Submitted: 2026-08-11
Updated: 2026-08-13
Code: https://github.com/stevietran/hvac_cdqerl
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: This paper proposes CQD-ERL, a contextual quality-diversity evolutionary reinforcement-learning controller for the supervisory control of a tropical, water-cooled chiller plant and its associated air
Terminology
Summary
This paper proposes CQD-ERL, a contextual quality-diversity evolutionary reinforcement-learning controller for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. Rather than converging to a single scalarised policy, the controller maintains a product archive of specialised policies indexed jointly by a data-driven operating context, a cluster of daily weather and load regime, and a context-invariant behaviour descriptor, filled by a gradient-free evolutionary operator and a soft-actor-critic policy-gradient operator that share one replay buffer. Every action is filtered through a deterministic safety shield before execution. The controller is trained on a two-tier reduced-order environment representing the latent load, cooling-tower approach and humidity constraints of a Singapore commercial building, and is evaluated over a full annual backtest against an ASHRAE Guideline 36 baseline.
The archive is a product structure: one behaviour-indexed sub-archive of twelve centroids is maintained per context cell, so a genome is scored inside one context and located there by behaviour. Context is defined empirically via principal-component projection of four daily features (mean cooling load, latent-load fraction, peak wet-bulb temperature, daily solar irradiation) onto two components explaining 77.1% of variance, then k-means clustering into sixteen cells plus two fixed sentinel centroids, giving eighteen context cells total. The behaviour descriptor is read directly from the policy by feeding a fixed synthetic probe battery to the actor and summarising emitted actions: regression slopes of chilled-water setpoint and tower-fan speed against outdoor temperature and humidity, plus the fraction of unoccupied probe points requesting the plant on, split below and above a 29°C drift threshold. Fitness is normalised against a precomputed Guideline 36 return for the same calendar day before insertion, collapsing between-context return variance to effectively zero.
The archive is improved each generation by two operators in fixed proportion: a gradient-free iso-line-directional evolutionary operator over two random elites, and a soft actor-critic policy-gradient operator that ascends an entropy-regularised value using twin critics trained on the shared replay buffer. Synergistic coupling includes periodic insertion of the running actor into the archive, critics learning from diverse elite rollouts, and warm-starting new actors by behaviour-cloning the highest-fitness elite. Every candidate action passes through a deterministic safety shield enforcing relative-humidity and dew-point ceilings, actuator ranges, ramp limits and chiller minimum on-off times, producing zero relative-humidity violations across all evaluations.
The environment is a two-tier reduced-order model: a millisecond-cost differential-algebraic system with a zone and air-side block updating on 5-minute physics sub-steps and a plant and water-side block updating on 20-minute control steps, calibrated toward Singapore conditions with a 0.71 sensible heat ratio, 25–26°C wet-bulb temperature, and SS 553 humidity limits treated as hard safety constraints. Training addresses partial observability through a warm-start bank of thermal-mass states stratified by day of week, since Monday's thermal mass carries the preceding weekend's heat.
Results over five seeds and a full-year backtest show the archive fills completely in every seed, reaching 50% coverage after 18,761 environment steps, 90% after 56,371 ± 9,833, and unity after 126,889 ± 15,324 steps. The quality-diversity score converged to 2190.70 ± 0.36 (coefficient of variation 0.017%), and all 1,080 stored elites across five seeds outperform Guideline 36 on their own context, with the best elite reaching F̃ = 0.404 ± 0.004 and the archive mean 0.1421 ± 0.0017. The context factor explains 99.43 ± 0.26% of between-cell fitness variance, while the behaviour factor explains only 0.079 ± 0.044%, confirming the behaviour axis is nearly orthogonal to quality.
Both learners reduce annual energy by statistically indistinguishable margins: 3.40 ± 0.43% (95% CI 2.86–3.94) for the archive and 3.67 ± 0.49% (95% CI 3.07–4.27) for the single SAC policy (Welch p = 0.383). The evolutionary component buys a 4.95× faster dispatch (1.29 ms versus 6.39 ms per control step), part-load efficiency, and reproducibility: annual chiller starts vary by 1.0% across CQD-ERL seeds and 18.9% across SAC seeds, a variance ratio of 272. At matched evaporator load, CQD-ERL improves on Guideline 36 in every load band, with the largest margins in the 100–400 RT part-load range: 10.4 to 17.8% at matched duty, versus 5.7 to 10.7% for SAC, making the archive 5.5 and 11.9 percentage points better at 100–200 RT and 200–300 RT respectively. The converged strategy runs chilled-water supply at 8.89°C against 7.68°C for G36 (+1.21 K), with 47% less setpoint modulation, 2.88 chillers staged against 2.08, and 11% lower tower-fan speed.
Saving by outdoor wet-bulb quartile falls monotonically and reverses sign in the top quartile: −9.50%, −6.36%, −2.97%, and +0.41% across quartiles, with plant-on fraction rising from 0.458 to 0.760. In the quartile carrying 37% of annual energy, both principal degrees of freedom are saturated, locating residual opportunity in load shifting and storage rather than better setpoints. Across five seeds, the relative-humidity violation rate was exactly zero, safety-shield correction magnitude was exactly zero, and the fallback hierarchy was never entered, with 100% primary-tier service. Occupied comfort was preserved (0.36 ± 0.32% of steps outside the PMV deadband, p = 0.208). All five seeds made exactly 45 context switches with a median dwell of 5.5 days, 83.3% of context cells and all behaviour niches were used, and dispatch entropy reached 74.7% of uniform.
Policy-gradient offspring exceed iso-line-directional offspring in mean normalised fitness by 0.117, or 82% of the mean elite fitness, while soft actor-critic actor injections at 0.62% of rollouts carry the highest mean fitness of any source. Principal-component analysis of the 216 stored 8,838-dimensional genomes shows the leading ten components capture a median 94.3% of elite variance, confirming the elite-hypervolume structure the iso-line operator exploits. The 3.40% aggregate saving sits far below the 20–70% of early deep reinforcement learning against unspecified baselines and below the 11–15% of field-validated model predictive control, explained by the Guideline 36 baseline, the absence of air-side economising at a 25°C wet bulb, and the un-priced flat tariff. The contribution is architectural: this is the first archive-based quality-diversity evolutionary reinforcement-learning controller applied to a building energy system, dispatched over a full year with zero violations, zero shield corrections and zero fallbacks at rule-based inference cost.
Improvements for AI systems
Improvements to AI systems:
-
Product-archive reinforcement learning with context-conditioned behavior descriptors: Replace single-objective RL with a multi-dimensional archive indexed by both operating context (e.g., weather/load clusters) and behavior (e.g., control policy signatures). This enables maintaining a diverse set of specialized policies rather than one averaged policy, improving robustness across heterogeneous conditions.
-
Hybrid evolutionary + policy-gradient optimization with shared replay buffer: Combine gradient-free evolutionary operators (iso-line-directional) with actor-critic policy gradients (soft actor-critic) in fixed proportion, sharing one replay buffer. This yields faster dispatch (4.95×), lower variance across seeds (272× reduction in chiller start variability), and better part-load efficiency without sacrificing optimality.
-
Deterministic safety shield as a hard filter on all actions: Enforce physical constraints (humidity ceilings, actuator limits, ramp rates, minimum on/off times) before execution, guaranteeing zero safety violations and zero corrective interventions. This makes RL safe for real-world deployment in safety-critical control.
-
Context-normalized fitness scoring: Normalize fitness against a precomputed baseline (e.g., Guideline 36) for the same calendar day before archive insertion, collapsing between-context return variance to near zero. This isolates quality from context, enabling fair comparison and faster convergence.
-
Warm-start bank of thermal-mass states stratified by day of week: Address partial observability by initializing hidden states (e.g., Monday’s thermal mass carrying weekend heat) from a precomputed bank, improving training stability and real-world transfer.
-
Principal-component-based context clustering: Reduce high-dimensional daily features (load, humidity, solar) to two principal components explaining 77.1% variance, then k-means cluster into 18 cells. This provides a compact, interpretable context space for indexing policies.
-
Behavior descriptor read directly from policy via synthetic probe battery: Summarize policy behavior by feeding fixed synthetic inputs and extracting regression slopes and activation fractions. This yields a low-dimensional, context-invariant behavior descriptor that is nearly orthogonal to quality (0.079% variance explained), enabling effective diversity search.
-
Elite-hypervolume exploitation via iso-line-directional operator: Use the leading principal components of stored elite genomes (capturing 94.3% variance) to guide mutation directions, improving search efficiency in high-dimensional policy spaces.
What the improved AI system can do:
-
Achieve near-zero safety violations (0% relative-humidity violations, 0 shield corrections) while maintaining comfort (0.36% PMV deadband excursions) in real-world building control.
-
Reduce annual energy consumption by 3.4–3.7% against a strong ASHRAE Guideline 36 baseline, with 10–18% savings in part-load ranges (100–400 RT), outperforming single-policy RL by 5.5–11.9 percentage points at matched duty.
-
Dispatch control actions in 1.29 ms (rule-based inference cost), enabling real-time deployment on edge hardware.
-
Maintain reproducible performance across seeds (coefficient of variation 0.017% in quality-diversity score), with chiller start variance reduced by 272× compared to standard SAC.
-
Adapt to diverse operating contexts (18 weather/load cells) with 100% archive coverage, automatically switching policies (median dwell 5.5 days) to match changing conditions.
-
Provide a diverse portfolio of specialized policies for each context-behavior niche, enabling operator selection for specific objectives (e.g., energy vs. comfort) without retraining.
-
Generalize to other complex control tasks (e.g., data center cooling, industrial HVAC, power grid management) where safety, diversity, and fast inference are critical.
Abstract
This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. Rather than converging to a single scalarised policy, the controller maintains a product archive of specialised policies indexed jointly by a data- driven operating context, a cluster of daily weather and load regime, and a context-invariant behaviour descriptor, filled by a gradient-free evolutionary operator and a soft-actor-critic policy-gradient operator that share one replay buffer. Every action is filtered through a deterministic safety shield before execution. The controller is trained on a two-tier reduced-order environment representing the latent load, cooling-tower approach and humidity constraints of a Singapore commercial building, and is evaluated over a full annual backtest against an ASHRAE Guideline 36 baseline.
Sources
- Multi-zone HVAC Control with Model-Based Deep Reinforcement Learning
- Data Center Cooling System Optimization Using Offline Reinforcement Learning
- Discovering the Elite Hypervolume by Leveraging Interspecies Correlation
- Quality Diversity for Multi-task Optimization
- Combining Evolution and Deep Reinforcement Learning for Policy Search: a Survey
- QDax: A Library for Quality-Diversity and Population-based Algorithms with Hardware Acceleration
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks