Surrogate Modeling of Interconnector Flows: A Machine Learning Alternative to Full-Scale Power System Simulations with Application to Cross-Border Electricity Exchange
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Surrogate Modeling of Interconnector Flows".
Dev: This paper proposes a machine-learning (ML) surrogate framework designed to generate synthetic, interconnector-level flow time series from readily available nodal data (demand and renewable generation).
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're looking at this paper now, "Surrogate Modeling of Interconnector Flows: A Machine Learning Alternative to Full-Scale Power System Simulations with Application to Cross-Border Electricity Exchange," and the authors are Robert Gaugl, Eloy Insunza, and José Portela. It seems like they’re tackling that big problem where we have to guess how much electricity will cross borders when renewable energy sources are changing so fast.
Dev: Exactly, Rosa; the title itself makes it clear they're trying to find a way around running those huge, slow full-scale Power System Optimization Models repeatedly just to see what happens with different climate scenarios. It’s about creating something much faster for decision-making and checking system responses.
Taro: I'm interested in how this relates to real-world autonomy; if we're looking at autonomous systems operating in a grid, knowing these cross-border flows accurately is vital because the energy supply isn't just local anymore.
Rosa: Well, the paper explains that they’ve developed a machine learning framework that takes simple nodal data—like demand and renewable generation—and uses it to synthesize interconnector flow time series instead of running the heavy simulations every single time.
Dev: That’s the core idea: mapping those available inputs directly to flows using ML so we don't have to repeatedly solve the full PSOM, which saves a massive amount of computational time, especially when you're testing many different scenarios.
Taro: It sounds like they’re trying to make something that can handle uncertainty better than just using old historical data because those patterns are changing with renewable penetration.
Rosa: Right, and the paper points out that traditional methods often simplify things by reusing historical time series for imports and exports, but this approach becomes inconsistent when renewable availability shifts significantly.
Dev: That inconsistency is a real headache; when you rely on fixed historical flows instead of letting the model predict them based on current conditions, your results won't match what’s actually happening in the system today.
Taro: So they are aiming to create flow profiles that are decision-relevant, meaning they actually reflect what the system *should* do under new renewable deployment plans.
Rosa: Precisely; their goal is to generate synthetic, decision-relevant flow profiles that can be used as fixed boundary conditions in reduced power system optimization models.
Dev: And they compare two specific ML families: a non-parametric k-nearest neighbors baseline and a feedforward neural network surrogate, which they call SQU.
Taro: I’m curious about the comparison; what makes the neural network approach better than the simpler KNN for capturing these complex interconnector dynamics?
Rosa: The paper shows that while KNN is a baseline, it doesn't generalize as well to unseen climate years and it tends to underperform when compared against scaled historical benchmarks in terms of predictive accuracy.
Title and authors: Dev: They find that the SQU models actually generalize more robustly than KNN, which means they perform better when we test them on data they haven't seen before, which is crucial for future planning.
Taro: That’s important because if a system behaves differently in an unseen climate year, we need our control systems to be prepared for that behavior too.
Rosa: The authors also introduce something quite interesting: a feasibility-aware training variant for the SQU model using a custom loss function. This function is designed to penalize any flow patterns that look physically unrealistic during the training process itself.
Dev: That loss function is key because it directly aims to reduce surrogate-induced infeasibilities, which the authors specifically call ENS, when those flows are later used as fixed inputs in reduced PSOMs.
Taro: So if the surrogate generates a flow pattern that would make the downstream model impossible to solve under certain conditions, this new training method tries to prevent that from happening at the source.
Rosa: Yes, and they demonstrate this is effective in Austria, where it eliminated those ENS issues when they ran reduced single-country simulations using these flow inputs.
Dev: It’s a direct attempt to improve the decision relevance of the surrogate by ensuring it produces flows that are physically plausible within the constraints of supply and demand.
Taro: That speaks to robustness; if we feed a model unrealistic data, we get unrealistic outputs, so making sure the input is realistic is a necessary step for any reliable AI deployment in this field.
Rosa: Moving on to how they test these models, they constructed feature vectors that include both "Full" sets—all demand and renewable generation profiles—and smaller "Selected" sets determined by feature importance analysis.
Dev: That selection process shows they’re trying to find the most critical inputs, which makes the training process much more efficient because you aren't feeding the model unnecessary noise.
Taro: It seems like an automated way to figure out what truly drives interconnector flows, rather than just throwing everything at the model and hoping for the best.
Rosa: Indeed, and they compare how these different feature sets perform across various configurations—full data versus selected data—across different system structures.
Dev: Their comparison results show that for systems like Austria and Germany, the standard KNN variants fall short compared to SQU models, with test R2 values hovering around zero point seven zero to zero point seven two and NMAE in the range of zero point three seven to zero point four one for KNN alone.
Taro: So the neural network approach has a clear advantage even when you only look at specific inputs, which suggests a stronger underlying mapping capability for these physical relationships.
Rosa: That's what they found; SQU models consistently outperform KNN variants across various configurations and feature sets, showing better predictive accuracy in several cases.
Title and authors: Dev: They also highlighted that for Spain, the SQU models showed a substantial improvement in explained variance, achieving test R2 values of zero point seven five with much lower NMAE around zero point three one to zero point three three when compared to KNN results there.
Taro: That difference in performance between countries is interesting; it means the model isn't just one-size-fits-all and adapts its modeling approach based on the specific system characteristics of that region.
Rosa: And we can’t forget the computational aspect, because they showed that running a single-country formulation with these ML flow surrogates is much faster than generating the full European model, leading to speedups of up to around five hundred times.
Dev: A five hundred times speedup is significant for operational tasks; it means we can run many more iterations or check more scenario combinations in a fraction of the time that it would take to solve the full system.
Taro: That computational saving is where this research really lands; it moves the problem from being intractable for real-time planning to being manageable.
Rosa: So, to wrap up these improvements, this paper suggests using SQU models because they offer a more consistent approximation of interconnector flows across different countries and climate years compared to the KNN baselines.
Dev: The feasibility-aware training variant is presented as the key addition for making sure that when we use these ML-generated flows as fixed inputs in those reduced PSOMs, we minimize those problematic surrogate-induced infeasibilities.
Taro: I think what this research really contributes is a methodology for generating synthetic data that respects the physical constraints of the power system from the start, which is a very practical way to improve autonomy in complex energy grids.
Rosa: So, to conclude on "Surrogate Modeling of Interconnector Flows: A Machine Learning Alternative to Full-Scale Power System Simulations with Application to Cross-Border Electricity Exchange," it provides a robust ML alternative for generating flows from nodal data, and the SQU models are shown to be the most consistent performers across different system structures.
Dev: The major implication is enabling faster scenario screening and more decision-relevant cross-border flow profiles, which is vital for planning under uncertainty in high renewable penetration environments.
Taro: I think we should focus on increasing spatial granularity next because while this framework works well for single nodes, scaling it up to capture more detailed spatial interactions would open up even more potential applications.
Rosa: And we'll definitely keep an eye on the work on probabilistic surrogates; moving from deterministic predictions to probabilistic ones will give us a much better measure of the risk involved in these flow predictions.
Dev: It sounds like this paper lays a solid foundation for using AI to quickly check system responses under different renewable penetration levels, provided we can manage the latency and ensure the failure modes don't creep in during implementation.
The paper's summary: Rosa: So, to recap, this paper proposes using machine learning surrogates to quickly generate interconnector flow time series from basic nodal data instead of running those heavy full-scale power system simulations every time we need a scenario tested.
Dev: That’s right; the core idea is replacing the slow PSOM runs with these fast ML approximations so we can get results much quicker for decision-making loops.
Taro: It really moves the problem from being computationally locked to being something that can be explored more frequently when things in the real world are changing rapidly.
Rosa: Exactly, and this isn't just about speed; it’s about consistency, because they are training these models on European data to ensure the synthetic flows actually make sense for cross-border exchanges.
Dev: And they did that by comparing two models: a simpler k-nearest neighbors approach and a more complex neural network surrogate called SQU.
Taro: I saw that the SQU model was showing better generalization capabilities, which means it handles those unseen climate years much better than the KNN baseline does.
Rosa: That's significant because if our planning models only work reliably on historical data, we’re stuck; this suggests a way for AI to keep making good predictions even when the conditions shift unexpectedly.
Dev: The authors also introduced a specific training trick with a custom loss function that actively penalizes physically impossible flow patterns during the learning process itself.
Taro: That feasibility constraint is what I find most compelling; it’s an attempt to build physical reality directly into the AI's training so it doesn't generate outputs that would be nonsensical in an actual power grid.
Rosa: It sounds like a very smart way to ensure the AI isn't just predicting numbers, but is learning what physically *can* happen in a power system.
Dev: From my side, it addresses the latency issue because if we can generate these boundary conditions so fast, our reduced optimization models can run much faster, which directly improves our control loop rate.
Taro: If we get this speedup and better constraint enforcement, imagine how much more nuanced and responsive our autonomous systems could be when dealing with unpredictable energy flows across borders.
Rosa: It really opens up possibilities for scenario screening on a massive scale; instead of testing one climate year at a time, you could test thousands in seconds.
Dev: That level of rapid iteration is what we need to stress-test the robustness of our control strategies under various extreme conditions.
Taro: So this work suggests that AI can act as a powerful tool for rapidly prototyping and validating complex, high-stakes planning scenarios that would otherwise be too slow to test manually.
Rosa: And while they show great results for specific countries like Austria and Germany, the paper does flag that scaling this up to capture more detailed spatial interactions is the next challenge.
Dev: That's fair; they focused on single-country runs initially, but getting it to work across a whole continent with fine spatial resolution would be the next big hurdle for deployment.
Taro: I think we should definitely keep an eye on those probabilistic surrogates mentioned in their conclusion because knowing the uncertainty around these flow predictions is as important as just having a single best guess.
Rosa: Absolutely, and this moves us toward a system where AI not only predicts what will happen but also tells us how confident it is in that prediction.
The paper's improvements: Rosa: So, we're looking at how this research actually improves things by focusing on what they suggested for future use and what that means for real systems.
Dev: It seems the main improvement is moving toward using these ML flow surrogates as actual fixed boundary conditions in reduced optimization models, which really cuts down on the computational load.
Taro: I think the implication here is that autonomous systems won't just be reacting to immediate local conditions, but they can plan across borders with much higher fidelity because they’re using more accurate flow data.
Rosa: Exactly; it means we can run these complex planning scenarios much faster and check if a proposed action will have ripple effects across different countries before committing to them.
Dev: And the feasibility-aware training component is a huge part of that improvement because it ensures the data we feed into the model is physically realistic, which minimizes those nasty surprises in the downstream optimization runs.
Taro: That’s what I like; if we can build AI systems that respect physical laws from the start during training, then when they encounter unexpected world events, their response won't be based on nonsensical assumptions.
Rosa: It really speaks to building more trustworthy autonomous agents; they won't make decisions based on flows that defy basic energy physics because the model was trained to avoid those impossibilities.
Dev: And I’m thinking about the latency aspect too; if the surrogate generates these flows quickly, it fits much better into a tight control loop where we need rapid feedback without waiting hours for a full simulation.
Taro: That speed and safety combination is what makes this applicable to high-stakes autonomous operations, where delays or physical inconsistencies can lead to serious failures.
Rosa: It’s about creating a fast, safe bridge between raw data and complex system planning that isn't bogged down by the massive computational requirements of traditional simulations.
Dev: So the next step seems to be scaling this up spatially; right now it works well for single nodes, but expanding it to capture detailed interactions across a whole continent is where we’ll see its full impact.
Taro: I agree; moving toward those larger scales will give autonomy researchers the tools they need to model interconnected, complex environments in much more realistic ways.
Rosa: And keep an eye on those probabilistic models, because knowing the uncertainty of these flow predictions will give us a much better measure of risk when deploying these AI-driven planning tools.
Conclusion: Rosa: So, to wrap things up on "Surrogate Modeling of Interconnector Flows: A Machine Learning Alternative to Full-Scale Power System Simulations with Application to Cross-Border Electricity Exchange," we've seen how machine learning can create very fast and physically consistent models for predicting electricity flows.
Dev: It really boils down to using SQU models trained with feasibility constraints to generate accurate boundary conditions for reduced optimization models, which drastically improves the speed of our planning loops.
Taro: I think the real impact is that we are moving toward a future where autonomous systems can plan across complex, interconnected energy markets without getting stuck waiting for massive simulation outputs.
Rosa: It's exciting because this technology allows us to test thousands of different climate and policy scenarios in seconds, which is exactly what we need for robust planning.
Dev: And the latency reduction from using these surrogates instead of full PSOM runs means we can actually integrate these flow predictions into real-time control systems much more effectively.
Taro: If autonomous systems can operate with this level of fast, reliable cross-border data, we open up entirely new ways to manage distributed energy resources and respond intelligently to sudden supply shocks.
Rosa: It’s a significant step toward making energy management proactive rather than reactive when things go wrong across different regions.
Dev: And while the authors did flag that scaling this up spatially is the next big engineering challenge, I think solving that will be key for wide deployment in real-world scenarios.
Taro: That's where autonomy research fits in; we need these tools to model those large, interconnected systems realistically when things go wrong in a way that isn't just a local glitch.
Rosa: So, this paper shows us a solid path forward for building more intelligent and responsive energy management AI.
Dev: We definitely need to keep pushing on the spatial resolution aspect; getting that granular is where we’ll see if this works reliably outside of the controlled lab environment.
Institute of Electricity Economics and Energy Innovation (IEE) · Research Center ENERGETIC · Institute for Research in Technology (IIT), ICAI School of Engineering, Universidad Pontificia Comillas
eess.SY, cs.SY
Submitted: 2026-06-02
Updated: 2026-09-30
Comments: Accepted manuscript. Revised following peer review. Published in Applied Energy
Journal ref: Applied Energy 427 (2027) 128946
DOI: 10.1016/j.apenergy.2026.128946
Code: https://github.com/pypsa/pypsa-eur
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 77/100
The gist: This paper proposes a machine-learning (ML) surrogate framework designed to generate synthetic, interconnector-level flow time series from readily available nodal data (demand and renewable
Key concepts
- Surrogate Modeling
- This involves using machine learning models to quickly approximate the results of complex, slow full-scale power system simulations. Instead of running heavy simulations repeatedly, these fast models generate synthetic data based on simpler inputs like demand and renewable generation.
- SQU Model
- The SQU model is a feedforward neural network surrogate used in the paper. It was shown to generalize more robustly than the k-nearest neighbors baseline, meaning it performs better when tested on unseen climate years and different system structures.
- Feasibility-Aware Training
- This is a training technique using a custom loss function that penalizes flow patterns during learning if they appear physically unrealistic. This prevents the surrogate from generating flows that would make the downstream power system optimization model impossible to solve, ensuring physical plausibility.
- Interconnector Flows
- These refer to the amount of electricity crossing borders between different power systems. Accurately predicting these flows is vital for planning and decision-making in cross-border electricity exchanges, especially when renewable energy sources are changing rapidly.
Terminology
Summary
This paper proposes a machine-learning (ML) surrogate framework designed to generate synthetic, interconnector-level flow time series from readily available nodal data (demand and renewable generation). This approach offers a computationally efficient alternative to running full-scale power system simulations, enabling decision-relevant cross-border flow profiles to be used as fixed boundary conditions in reduced power system optimization models (PSOMs), which is crucial for analyzing systems with high renewable penetration.
Motivation and Problem Statement
Cross-border electricity exchanges are critical for operating and planning highly renewable power systems, but traditional methods often reduce spatial granularity or prescribe exchanges exogenously using historical time series. This leads to inconsistencies as renewable penetration changes the magnitude and timing of flows. The paper addresses the need for decision-relevant surrogate modeling,
training on flows generated by a pan-European PSOM to create consistent flow profiles that are directly transferable to prospective scenario studies where renewable deployment and demand patterns are defined exogenously.
Methodology: Surrogate Framework
The proposed framework maps node-level inputs (demand and renewable generation) to interconnector flows using a machine learning surrogate, aiming to approximate the full PSOM output. The methodology involves:
-
Using the open-source LEGO PSOM to generate training data for the ML models and to compute reference results.
-
Training two surrogate families: a non-parametric k-nearest neighbors (KNN) baseline and a feedforward neural-network surrogate (SQU).
-
Constructing feature vectors, including
Full
sets (all demand and renewable generation profiles) andSelected
sets, which are determined by feature importance analysis to reduce dimensionality. -
Implementing a novel feasibility-aware training variant for the SQU model via a custom loss function that
penalizes physically unrealistic flow patterns.
Surrogate Model Comparison
The paper compares KNN and SQU models across various configurations (FULL, SELECTED, OPTIMIZED) using both full and reduced feature sets. Key findings include:
The SQU models generalize more robustly than KNN to unseen climate years and substantially improve upon scaled historical benchmarks in terms of predictive accuracy.
For Austria (AT) and Germany (DE), the standard KNN variants... remain clearly below the SQU family, with test R2 values around 0.70 – 0.72 and test NMAE in the range of 0.37 – 0.41.
For Spain (ES)... the SQU models improve explained variance substantially, achieving test R2 values of 0.75 with markedly lower test NMAE (0.31 – 0.33).
Feasibility-Aware Training and Decision Relevance
A novel custom loss function, denoted as Lcustom(θ), is introduced for the SQU model to discourage interconnector flows that are physically impossible.
This penalty activates when the predicted net export exceeds a precomputed residual bound R max t, which represents an upper bound on net export given demand and maximum available nodal supply. The goal is to reduce surrogate-induced infeasibilities (e.g., ENS) when flows are used as fixed inputs in reduced PSOMs.
This mechanism is shown to be effective in Austria, where it eliminates ENS in the reduced single-country runs,
indicating that feasibility-aware training can improve decision relevance.
Impact on Power System Optimization
The trained ML-generated flow time series are imposed as fixed boundary conditions
in a simplified PSOM with spatial resolution restricted to a single node. The evaluation assesses decision relevance by comparing deviations in key outputs (objective value, trade balances, and dispatch indicators) between reduced- and full-scale PSOM runs.
Results show that SQU surrogates generally provide better consistency across countries and climate years compared to KNN variants. Furthermore, the framework offers significant computational savings: PSOM model generation becomes much faster because the single-country formulation contains far fewer components than the full European model,
leading to speedups of up to ∼500×.
The SQU CUSTOM LOSS formulation is noted as being beneficial in some settings but can induce a systematic shift in dispatch patterns for Germany, moving it away from the benchmark trajectory.
Conclusion and Outlook
The paper concludes that ML-based surrogates provide consistently accurate and robust approximation of cross-border interconnector flows across diverse system structures.
The SQU models are identified as the most consistent performers across countries and climate years,
demonstrating a stronger ability to generalize nonlinear relationships than KNN baselines. The framework is valuable because it allows for fast evaluation of multiple climate years via single-country PSOM runs, and further work should focus on increasing spatial granularity and developing probabilistic surrogates. Furthermore, feasibility-aware training is most effective for eliminating surrogate-induced infeasibilities when the benchmark itself is feasible.
Key Enumerated Contributions:
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this paper, which proposes a machine learning (ML) surrogate framework for predicting cross-border electricity flows in power systems.
The primary contribution is replacing computationally expensive full-scale Power System Optimization Models (PSOMs) with fast ML surrogates that map node-level inputs to interconnector flows. The second major contribution is the introduction of a custom loss function to enforce physical feasibility constraints during training, leading to more reliable boundary conditions for downstream reduced PSOMs.
Here are the specific improvements and capabilities this research enables in AI systems:
)Specific Improvements and Enhanced AI Capabilities:
AI-Driven Scenario Screening (Fast & Consistent):
Predict interconnector flows instantly using ML surrogates trained on historical data (or climate year inputs). This allows for the rapid evaluation of thousands of potential cross-border scenarios—such as different renewable penetration levels, new interconnection policies, or varying demand patterns—in seconds rather than hours.
Reduced PSOM Optimization (Tractable & Decision-Relevant):
The ML flow surrogates serve as fixed boundary conditions in reduced single-node DC Optimal Power Flow (DC-OPF) models. This enables the AI system to perform optimization tasks that are otherwise computationally prohibitive, such as:
-
Determining optimal dispatch schedules for local generation and storage units.
-
Calculating accurate trade balances and operational costs for a specific country under a given flow constraint.
-
Assessing the impact of cross-border constraints on local system performance (e.g., energy not supplied (ENS) indicators).
Robustness Against Climate Uncertainty (Generalization):
The use of advanced neural network architectures, specifically the Sequential Neural Network Surrogate (SQU), demonstrates superior generalization compared to simpler models like K-Nearest Neighbors (KNN). This means the AI system can provide reliable flow predictions for unseen climate years (like 1995 or 2008) that are not in the training data, ensuring that planning decisions remain robust even as renewable availability and demand patterns shift.
Feasibility-Aware Constraint Learning (Safety & Reliability):
The custom loss function penalizes physically impossible flow patterns during training (e.g., predicted exports exceeding available domestic supply). This directly improves the reliability of the surrogate when used in real optimization settings, as it steers the model away from generating boundary conditions that would lead to artificial system infeasibilities (ENS) in the downstream reduced PSOM.
Automated Feature Engineering (Data-Driven):
The study identifies and compares Full
feature sets (all demand/generation data) versus Selected
feature sets based on permutation importance analysis. This provides an automated method for AI systems to identify the most critical inputs driving cross-border flows, allowing the system to operate efficiently even when dealing with high-dimensional datasets.
Quantifiable Model Risk Assessment (Fidelity Metrics):
The framework provides clear metrics (NMAE and R2) to quantify the predictive accuracy and variance explained for different model configurations (KNN vs. SQU). This allows researchers to make informed trade-offs between model complexity, development cost, and predictive performance, ensuring that the chosen AI surrogate meets specific engineering requirements for a given application.
In summary, this research improves AI systems by moving beyond simple forecasting to creating a closed-loop system: predicting complex physical interactions (flows) accurately and safely (feasibility-aware) so that high-level planning models (PSOMs) can run extremely fast, providing decision-relevant insights under uncertainty.
Abstract
Cross-border electricity exchanges are crucial for operating and planning highly renewable power systems. Many studies reduce spatial granularity to keep models tractable and prescribe cross-border exchanges exogenously, often by reusing historical import/export time series. Such assumptions become inconsistent as renewable penetration changes the magnitude and timing of flows. The core contribution of this work is a machine-learning (ML) surrogate framework that maps nodal time-series data (e.g., hourly demand and renewable generation) to synthetic interconnector-level flow time series. The resulting flows serve as fixed boundary conditions in reduced power system optimization models (PSOMs). This addresses studies focused on a subset of an interconnected system, for which repeatedly solving the full model across multiple climate years or scenarios is computationally expensive. We demonstrate the framework on a pan-European single-node-per-country DC optimal power flow model. We benchmark k-nearest neighbors, ridge regression, random forests, Extra Trees, gradient boosting, XGBoost, and feedforward neural-network (SQU) surrogates. As an additional contribution, we introduce a novel custom neural-network loss that softly penalizes physically implausible flow patterns. This soft penalty is not a guarantee of full power-flow feasibility. Across benchmarks, the SQU variants provide the most consistent out-of-sample and downstream optimization performance while substantially outperforming scaled historical profiles. The ML-generated flows closely reproduce country-level results of the full European model while greatly reducing computation time. The proposed framework enables scenario-consistent reduced PSOM studies across climate years after a single full-model run, eliminating repeated pan-European simulations and supporting more robust climate-year analyses under fixed system assumptions.
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation