Skill-Based AI Agents for Power-System Studies
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Skill-Based AI Agents for Power-System Studies".
Rosa: This paper describes a skill-based agentic framework for power-system studies using Model Context Protocol (MCP)-connected engineering tools,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, Dev, we're looking at this paper on "Skill-Based AI Agents for Power-System Studies." Basically, they've put together a framework that uses these agentic systems to work with real engineering tools like PSS®E to speed up those power system simulations.
Dev: Right, Rosa? It seems the core idea is using an orchestration agent that takes a study objective and breaks it down into smaller tasks for specialized agents to handle, which then use reusable skill files to manage the procedures.
Taro: That sounds like they’re trying to give the AI a structured way to actually *do* the work rather than just chatting about power systems.
Rosa: Exactly, and what matters is that they connect these agents directly to engineering tools through something called Model Context Protocol, or MCP. This MCP server acts as the controlled interface allowing the agents to interact with PSS®E for things like running power-flow analysis or dynamic simulations.
Dev: I see how it matters for loop rates and latency because having a structured way for the AI to call those specific functions means we can potentially get much tighter control over how fast and reliably those simulations run.
Taro: When you talk about those specific functions, what's the biggest benefit they claim this architecture offers compared to just running PSS®E manually?
Rosa: Well, the abstract states that these agentic systems can greatly accelerate the power system dynamic simulation process for transmission planning studies by leveraging industry-grade simulation platforms. This means we could get those results much faster for planning purposes.
Dev: Faster execution is critical, but I'm wondering about the reliability of this whole setup when things go sideways in a real-world scenario. How robust are these skill files against unexpected inputs or tool errors?
Taro: That’s a big question, because if the system misbehaves when the world gets messy, what happens to that acceleration they're promising?
Rosa: The paper mentions that the reusable skill files encode failure-handling rules and validation checks within them, which is supposed to prevent them from improvising unsupported values or actions when inputs are missing or inconsistent.
Dev: That sounds like a necessary safeguard for any system interfacing with complex engineering software; we don't want it making things up.
Taro: I'm thinking about what happens when the world misbehaves, like during a major disturbance event. Does this skill-based approach allow the agents to handle those unexpected situations intelligently, or is it just limited to the defined procedures?
Rosa: They are designed to handle failure-handling rules specifically for those situations, suggesting an attempt at intelligent response rather than just stopping.
Dev: From my end, if the MCP server provides structured tool interfaces and error propagation, that should help manage the complexity of the PSS®E interactions without creating chaotic loops or unpredictable delays in execution.
Taro: So they're focused on making sure that even when things go wrong during a dynamic simulation setup, the agent doesn't just crash?
Paper summary: Rosa: That’s exactly what they seem to be targeting, ensuring that the agent reports missing or corrupted inputs instead of trying to guess what the right value is.
Dev: And this whole thing is being tested across three representative study procedures: power-flow and case-analysis, dynamic simulation, and a play-in approach for model validation using synchrophasor or SCADA data.
Taro: Those three procedures cover a good range of use cases; testing them from steady state checks all the way up to comparing simulated responses against measured ones sounds like they're hitting all the necessary angles for real application.
Rosa: It really shows the versatility of this framework, moving beyond just one type of analysis and into validation methods.
Dev: I do think that linking dynamic simulation with model validation using data like synchrophasor measurements is where we see a lot of practical value, especially when we’re trying to ensure our models are actually accurate for real system behavior.
Taro: If the AI can compare its simulated responses against actual measured ones, that gives us a much stronger confidence in the results derived from these power-system studies.
Rosa: The paper's main argument is that this shift in approach allows engineers to focus more on scenario design and interpretation rather than spending all their time operating the specific tool interfaces themselves.
Dev: That sounds like a significant change for how we structure our work in transmission planning; moving away from direct tool manipulation toward high-level objective setting for the AI.
Taro: I think the long-term implication is that this could allow us to explore much more complex scenarios and analysis methods than we currently have the time or resources to execute manually.
Rosa: So, in simple terms, this paper describes an agentic framework that uses structured skills and an MCP connection to make power system dynamic simulations much faster for transmission planning studies.
Dev: And it's evaluating two different ways to build it—one using the OpenAI Agents SDK and another with a Claude Code command-line interface—to see which implementation path is more practical for different kinds of customization.
Taro: That comparison between the two platforms is interesting because it shows that there are different routes to building these systems, each with its own trade-offs regarding development effort.
Rosa: I think the paper's title, "Skill-Based AI Agents for Power-System Studies," really captures the essence of what they’ve built: using skills and agents to handle the specific domain knowledge required for these engineering tasks.
Dev: It seems like the authors are pointing toward a future where engineers collaborate with an AI system that handles the heavy lifting of simulation setup and result extraction.
Taro: If this holds up outside of a controlled lab environment, that's where I want to see it tested next; we need to know how long these systems can run reliably when they aren't being fed perfectly curated data from a dataset.
Rosa: That’s the question for the future, isn't it? We need those real-world deployment scenarios to see if this accelerated process translates into actual efficiency gains for transmission planning.
Conclusion: Rosa: So, this paper is about using skill-based AI agents to speed up power system studies by connecting them to tools like PSS®E through an MCP server.
Dev: I'm focused on how fast those simulations run and what happens when they encounter errors during the process.
Taro: I'm curious about the autonomy aspect, specifically what these agents do when things go wrong in a dynamic simulation environment.
Rosa: Thinking about that title, "Skill-Based AI Agents for Power-System Studies," it really boils down to giving an AI a structured way to handle complex engineering tasks.
Dev: I see how that structure helps manage the loop rates and latency issues we worry about when running these simulations manually.
Taro: And I want to know if those skill files give the AI enough autonomy to make smart decisions when the simulation doesn't go exactly as planned during a disturbance setup.
Rosa: The authors are demonstrating that this framework lets agents handle routine simulation setup and result extraction, freeing up engineers to focus on designing the scenarios themselves.
Dev: That sounds like it could really change our workflow by shifting our focus away from tedious tool operation toward higher-level planning and interpretation.
Taro: It feels like a big step toward systems that can manage the complexity of dynamic simulations without needing constant, minute control from a human operator.
Rosa: And considering the authors, they've shown how different implementation pathways exist, which suggests this approach is flexible enough for various development needs across different engineering teams.
Dev: That flexibility is important because it means we can tailor the setup to fit our specific requirements regarding tool interaction and error handling.
Taro: I wonder what the long-term autonomy looks like when we apply these agents to much more complex analyses, like contingency planning or oscillation studies.
PNNL
eess.SY, cs.SY
Submitted: 2026-09-30
Updated: 2026-10-01
Comments: submitted to 2027 IEEE PES Grid Edge Technologies Conference & Exposition
Code: https://github.com/Power-Agent/PowerMCP
Project page: https://openai.github.io/openai-agents-python
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 82/100
The gist: This paper describes a skill-based agentic framework for power-system studies using Model Context Protocol (MCP)-connected engineering tools, demonstrating that agentic systems can greatly accelerate
Key concepts
- Orchestration Agent
- This agent receives a high-level study goal, breaks it down into smaller steps, and assigns those steps to specialized task agents. It acts as the project manager, ensuring all necessary subtasks are completed in the correct sequence to achieve the overall objective.
- Skill Files
- These reusable files contain detailed procedural instructions for specific tasks. They define exactly what inputs are needed, the sequence of tool calls required, how to validate results, and rules for handling errors during execution.
- MCP Server
- This server acts as a controlled bridge between the AI agents and engineering software like PSS®E. It exposes complex software functions as standardized tools with structured inputs and outputs, allowing agents to interact with the simulation without needing direct system access.
- Agentic Framework
- This is a multi-agent architecture where different LLM agents collaborate. They work together by delegating specific parts of a complex problem to specialized agents, enabling them to perform sophisticated tasks like setup, execution, and result extraction.
Terminology
Summary
This paper describes a skill-based agentic framework for power-system studies using Model Context Protocol (MCP)-connected engineering tools, demonstrating that agentic systems can greatly accelerate power system dynamic simulation processes for transmission planning studies.
The core concept involves a skill-based multi-agent architecture that coordinates LLM agents to interact with deterministic engineering tools like Siemens PTI PSS®E.
The proposed framework uses an orchestration agent
that receives a study objective, decomposes it into subtasks, and delegates execution to specialized task agents. These task agents are supported by reusable skill files
which encode crucial procedural instructions, including required inputs, toolcall sequences, validation checks, failure-handling rules,
and reporting templates.
MCP servers provide the controlled interface to PSS®E functions such as case management, power-flow solution, dynamic-simulation setup, simulation execution, result extraction, and model validation.
The framework is implemented through two distinct pathways: one using the OpenAI Agents SDK and another using a Claude Code command-line interface (CLI).
The OpenAI Agents SDK implementation uses Python for agent orchestration
and provides flexible control over orchestration, tools, guardrails, tracing, and MCP interfaces.
The Claude Code CLI implementation acts as a primary session coordinator,
enabling agents to determine whether to use an available skill, delegate tasks to a PSS®E subagent via the MCP server, or execute local post-processing commands through its shell environment.
The architecture supports three representative study procedures: powerflow and case-analysis tasks, dynamic simulation, and play-in approach for model validation.
The first procedure addresses powerflow and case-analysis tasks,
including case loading, solution checking, data-quality review, and summary reporting.
The second covers dynamic simulation,
which involves disturbance setup, simulation execution, channel extraction, and numerical-quality checks.
The third applies the play-in approach for validating models using measurements like synchrophasor or SCADA data to compare simulated responses against measured ones.
The MCP server serves as the interface layer between LLM agents and engineering software.
A custom MCP server
is implemented in Python using the FastMCP framework to expose PSS®E functions as MCP-compliant tools with structured inputs and outputs.
This design ensures that agents do not directly manipulate PSS®E or the operating system, making the tool interface more modular, auditable, and reusable across different agent runtimes.
The server supports 21 PSS®E tools in the tested configuration.
The evaluation compared two foundation models—GPT5.5 via OpenAI Agents SDK and Claude Sonnet 4.6 via Claude Code CLI—on representative study tasks.
Both frontier-model-based implementations successfully executed representative study tasks,
showing qualitatively similar capability for the tested tasks.
A key implementation difference noted is that the Claude Code implementation required less custom software development
because it provided an integrated runtime, whereas the OpenAI Agents SDK required these capabilities to be be assembled and customized through Python code.
The results showed that agentic systems can improve efficiency and repeatability in power-system analysis.
The paper concludes that agentic approaches can improve the efficiency and repeatability of common power-system analysis tasks.
The authors suggest that this points toward a shift where agents handle routine simulation setup and result extraction,
allowing engineers to focus on scenario design and interpretation rather than tool operation.
Future work will expand MCP functions to support tasks like contingency analysis, batch dynamic simulations, model-parameter calibration, oscillation analysis, and automated reporting.
Future research will explore secure deployment strategies for production environments.
Future work includes evaluating on-premises and grid-edge deployment using locally hosted models
and exploring secure execution environments for applications involving confidential planning models and sensitive data,
addressing concerns regarding the use of provider-hosted frontier models in critical infrastructure applications.
Key components of the framework include:
-
An orchestration agent to decompose objectives into subtasks.
-
Specialized task agents guided by reusable skills encoding procedures and validation checks.
-
An MCP server to provide controlled access to PSS®E functions via structured interfaces (FastMCP).
-
Two implementation pathways: OpenAI Agents SDK (Python-based) and Claude Code CLI (skill/subagent-oriented).
-
A sandboxed Python environment for code and visualization agents to execute generated scripts separately from the PSS®E execution environment.
Key findings regarding implementation trade-offs:
: The Claude Code implementation was more practical when complex customization was not required,
requiring less custom software development.
: The OpenAI Agents SDK offers greater flexibility for customized orchestration and production-style applications.
The overall conclusion is that LLM agents can effectively orchestrate them through structured workflows informed by huma subject matter expertise
rather than replacing validated analysis tools.
Improvements for AI systems
Here are specific improvements to AI systems based on this research, detailing what the improved system can achieve:
-
The core capability of AI will shift from mere conversational assistance to acting as an autonomous, multi-step engineering workflow orchestrator for power-system studies.
-
AI agents can independently manage the entire end-to-end simulation lifecycle for complex power system scenarios, reducing the need for human experts to manually execute repetitive setup and extraction tasks.
-
The improved system can decompose high-level study objectives (e.g.,
Evaluate contingency impact on stability
) into a structured sequence of deterministic engineering tasks (power flow setup, dynamic simulation execution, result extraction), delegating these steps to specialized subagents. -
AI systems can leverage reusable
skills
that encode domain-specific procedural knowledge—such as the exact sequence of PSS®E tool calls for specific analysis types or the precise logic for play-in model validation checks—ensuring procedural consistency and accuracy across different studies. -
The system can interface with industry-grade, deterministic simulation platforms (like PSS®E) via a standardized Model Context Protocol (MCP) server, allowing LLMs to orchestrate these tools without directly manipulating complex engineering software or operating systems, ensuring traceability and auditability.
-
The AI can perform automated result post-processing and visualization: given raw simulation outputs (e.g., CSV files), the agent can automatically identify relevant data points (time series, specific channel values), generate necessary plotting scripts, apply domain-specific transformations (like frequency conversion from per unit to Hertz), and produce publication-ready figures.
-
The system can execute complex validation workflows autonomously: it can coordinate multiple data sources (measured event records, dynamic simulation outputs, calibrated/uncalibrated model responses) to run sophisticated play-in tests for power plants or inverter models and generate comparative analysis plots between simulated and measured responses.
-
The AI architecture offers flexibility in deployment:
Ease of prototyping and rapid iteration can be achieved using provider-hosted frontier models (like GPT-5.5), while a production environment can transition to secure, on-premises, or grid-edge deployments utilizing open-weight models like Claude Sonnet 4.6 or locally hosted NVIDIA NemoClaw systems for sensitive data control and offline operation.
- The system can integrate real-time data streams by connecting through APIs (e.g., PMU signature libraries), enabling proactive workflows where the agent can automatically screen disturbance records in real-time and trigger domain-specific studies (like oscillation analysis) as soon as a qualified event is detected.
Abstract
This paper describes a skill-based agentic framework for power-system studies using Model Context Protocol (MCP)-connected engineering tools. A custom MCP server was developed to expose Siemens PTI PSSE functions for power-flow analysis, dynamic simulation, result extraction, and model-validation workflows. Two implementation pathways built on a programmable OpenAI Agents software development kit (SDK) and a Claude Code command-line interface (CLI) were evaluated, both using reusable skills, subagents, MCP tools, data-repository connections, and local shell/Python execution. Both frontier-model-based implementations successfully executed representative study tasks. Success was evaluated based on task completion, output accuracy, and the need for human expert interventions. Results based on public datasets show that agentic systems can greatly accelerate power system dynamic simulation process for transmission planning studies leveraging industry-grade simulation platforms. This points toward a shift in transmission planning practice, where agentic systems could handle routine simulation setup and result extraction, allowing engineers to focus expert judgment on scenario design and interpretation rather than tool operation.
Sources
- PowerDAG: Supervisory Agentic AI System for Automating Distribution Grid Analysis
- Grid-Agent: An LLM-Powered Multi-Agent System for Power Grid Control
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation