5.3SYJun 30
A Shallow Recurrent Decoder for Dynamic State Estimation with a Limited Number of PMUs in Power SystemsAndrea Pomarico, Alberto Berizzi, J. Nathan Kutz
Dynamic State Estimation (DSE) will play a fundamental role in future power system operation by providing real-time estimates of the system state and enabling enhanced situational awareness. Existing DSE approaches are primarily based on Kalman filter variants or Machine Learning (ML) techniques. However, Kalman-based methods often suffer from high computational complexity, sensitivity to model inaccuracies, and performance degradation under strongly nonlinear operating conditions. Moreover, their effectiveness critically depends on the number and placement of measurements, since suboptimal PMU locations can reduce observability and even render state estimation infeasible. Machine learning approaches alleviate some of these limitations but typically require large amounts of training data and may struggle to generalize. To address these challenges, this paper proposes a SHallow REcurrent Decoder (SHRED) architecture for full-state reconstruction of power systems from sparse measurements. Unlike conventional model-based estimators, the proposed approach does not rely on an accurate physical model and is largely insensitive to PMU placement, making it particularly attractive for practical deployment in existing Wide Area Measurement Systems (WAMS). The method is validated on the IEEE 39-bus system under strongly nonlinear conditions, including short-circuit disturbances. The results demonstrate that SHRED can accurately reconstruct the complete system state using only a limited number of PMU measurements, consistently outperforming a state-of-the-art shallow decoder benchmark in sparse-measurement scenarios. Furthermore, the proposed framework exhibits strong robustness to measurement noise and maintains high reconstruction accuracy even under severe disturbances, highlighting its potential as a scalable and reliable alternative to conventional DSE techniques.
5.6SYJun 17
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State StudiesCostas Mylonas, Magda Foti, Andrea Pomarico et al.
Power system benchmarks usually evaluate numerical solvers, prediction models, or sequential controllers. These benchmarks are necessary, but they do not directly test whether a Large Language Model (LLM) agent can execute an engineering workflow: inspect a grid case, select tools, call simulators, screen contingencies, propose admissible mitigations, validate results, and produce an auditable evidence trail. This paper introduces PowerAgentBench-SS, a steady-state benchmark framework for evaluating tool-using agents in power system operation and planning studies. The benchmark exposes public case data, action constraints, a tool API, and a validation budget to an agent, while a hidden evaluator recomputes physical validity and scores the submitted report. We define the agent interface, tool contract, evidence log, and risk-sensitive metrics, including submitted recall, evidence-backed recall, found recall, false-safe penalties, severity regret, residual violation score, action cost, tool-use efficiency, and workflow diagnostics. To make the framework concrete, we instantiate the protocol in a reproducible DC thermal N-2 contingency-search pilot on deterministic IEEE 39-bus operating-point variants, with scripted baselines, an LLM JSON-command adapter, three locally hosted Ollama LLM agents, and one OpenAI API agent. The results show why solver-only or answer-only evaluation is insufficient: agents are distinguished not only by top-contingency discovery, but also by validation-budget use, explicit submission, type coercions, duplicate validations, evidence-backed reporting, and mitigation behavior.
3.3SYJun 18
PowerAgentBench-Dyn: A Benchmark for Agentic AI in Power System Dynamic StudiesQian Zhang, Andrea Pomarico, Costas Mylonas et al.
Large Language Model (LLM)-based agents are increasingly being used to automate multi-step engineering work flows by interacting with software tools, interpreting intermediate results, and autonomously planning subsequent actions. Power system dynamic studies represent a particularly promising yet largely unexplored application domain for these agents. Unlike static computational tasks, dynamic studies often require more time on model parameter calibration, engineering judgment, and decision making under constrained action spaces. This paper introduces PowerAgentBench-Dyn, a benchmark designed to evaluate Agentic AI systems on power system dynamic-analysis tasks. The benchmark targets problems that cannot be reduced to a single optimization or coding task, but instead require a type of reasoning, tool usage, and iterative experimentation routinely performed by experienced power system engineers. The proposed framework includes two initial benchmark tasks. The first, the Dynamic Model Quality Review Benchmark, evaluates agents' ability to validate and diagnose dynamic models based on model-quality compliance criteria specified by system operators. The second, the Dynamic Security Risk Screening Benchmark, assesses agents' capability to leverage semantic memory and a limited simulation budget to identify, rank, and analyze the most critical short-circuit contingencies from an unseen fault dataset, as well as propose and evaluate possible mitigation measures. For each task, we define the simulation environment, observation and action spaces, and evaluation metrics. The benchmark is reproducible in a metric-based sense: released cases and simulator settings define a deterministic evaluator, while stochastic agent behavior is assessed over repeated runs using success rates and other metrics. The benchmark supports the development of future Agentic AI for power system operation and planning.