2.9OCMar 15, 2016
Tight LP Approximations for the Optimal Power Flow ProblemSleiman Mhanna, Gregor Verbic, Archie Chapman
DC power flow approximations are ubiquitous in the electricity industry. However, these linear approximations fail to capture important physical aspects of power flow, such as the reactive power and voltage magnitude, which are crucial in many applications to ensure voltage stability and AC solution feasibility. This paper proposes two LP approximations of the AC optimal power flow problem, founded on tight polyhedral approximations of the SOC constraints, in the aim of retaining the good lower bounds of the SOCP relaxation and relishing the computational efficiency of LP solvers. The high accuracy of the two LP approximations is corroborated by rigorous computational evaluations on systems with up to 9241 buses and different operating conditions. The computational efficiency of the two proposed LP models is shown to be comparable to, if not better than, that of the SOCP models in most instances. This performance is ideal for MILP extensions of these LP models since MILP is computationally more efficient than MIQCP.
1.2SYJul 17, 2020
Probabilistic assessment of the impact of flexible loads under network tariffs in low voltage distribution networksDonald Azuatalam, Archie C. Chapman, Gregor Verbič
Given the historically static nature of low-voltage networks, distribution network companies do not possess tools for dealing with an increasingly variable demand due to the high penetration of distributed energy resources (DER). Within this context, this paper proposes a probabilistic framework for tariff design that minimises the impact of DER on network performance, stabilise network company revenue, and improves the equity of network costs allocation. To address the issue of the lack of customers' response, we also show how DER-specific tariffs can be complemented with an automated home energy management system (HEMS) that reduces peak demand while retaining the desired comfort level. The proposed framework comprises a nonparametric Bayesian model which statistically generates synthetic load and PV traces, a hot-water-use statistical model, a novel HEMS to schedule customers' controllable devices, and a probabilistic power-flow model. Test cases using both energy- and demand-based network tariffs show that flat tariffs with a peak demand component reduce the customers' cost, and alleviate network constraints. This demonstrates, first, the efficacy of the proposed tool for the development of tariffs that are beneficial for networks with a high DER penetration, and second, how customers' HEM systems can be part of the solution.
3.2ROMar 14, 2025
Training Directional Locomotion for Quadrupedal Low-Cost Robotic Systems via Deep Reinforcement LearningPeter Böhm, Archie C. Chapman, Pauline Pounds
In this work we present Deep Reinforcement Learning (DRL) training of directional locomotion for low-cost quadrupedal robots in the real world. In particular, we exploit randomization of heading that the robot must follow to foster exploration of action-state transitions most useful for learning both forward locomotion as well as course adjustments. Changing the heading in episode resets to current yaw plus a random value drawn from a normal distribution yields policies able to follow complex trajectories involving frequent turns in both directions as well as long straight-line stretches. By repeatedly changing the heading, this method keeps the robot moving within the training platform and thus reduces human involvement and need for manual resets during the training. Real world experiments on a custom-built, low-cost quadruped demonstrate the efficacy of our method with the robot successfully navigating all validation tests. When trained with other approaches, the robot only succeeds in forward locomotion test and fails when turning is required.
4.1LGMar 14, 2025
Low-cost Real-world Implementation of the Swing-up Pendulum for Deep Reinforcement Learning ExperimentsPeter Böhm, Pauline Pounds, Archie C. Chapman
Deep reinforcement learning (DRL) has had success in virtual and simulated domains, but due to key differences between simulated and real-world environments, DRL-trained policies have had limited success in real-world applications. To assist researchers to bridge the \textit{sim-to-real gap}, in this paper, we describe a low-cost physical inverted pendulum apparatus and software environment for exploring sim-to-real DRL methods. In particular, the design of our apparatus enables detailed examination of the delays that arise in physical systems when sensing, communicating, learning, inferring and actuating. Moreover, we wish to improve access to educational systems, so our apparatus uses readily available materials and parts to reduce cost and logistical barriers. Our design shows how commercial, off-the-shelf electronics and electromechanical and sensor systems, combined with common metal extrusions, dowel and 3D printed couplings provide a pathway for affordable physical DRL apparatus. The physical apparatus is complemented with a simulated environment implemented using a high-fidelity physics engine and OpenAI Gym interface.
1.0LGJun 11, 2019
Macro-action Multi-time scale Dynamic Programming for Energy Management in Buildings with Phase Change MaterialsZahra Rahimpour, Gregor Verbic, Archie C. Chapman
This paper focuses on energy management in buildings with phase change material (PCM), which is primarily used to improve thermal performance, but can also serve as an energy storage system. In this setting, optimal scheduling of an HVAC system is challenging because of the nonlinear and non-convex characteristics of the PCM, which makes solving the corresponding optimization problem using conventional optimization techniques impractical. Instead, we use dynamic programming (DP) to deal with the nonlinear nature of the PCM. To overcome DP's curse of dimensionality, this paper proposes a novel methodology to reduce the computational burden, while maintaining the quality of the solution. Specifically, the method incorporates approaches from sequential decision making in artificial intelligence, including macro actions and multi-time scale Markov decision processes, coupled with an underlying state-space approximation to reduce the state-space and action-space size. The performance of the method is demonstrated on an energy management problem for a typical residential building located in Sydney, Australia. The results demonstrate that the proposed method performs well with a computational speed-up of up to 12,900 times compared to the direct application of DP.
1.2SYApr 13, 2019
A Novel Probabilistic Framework to Study the Impact of PV-battery Systems on Low-Voltage Distribution NetworksYiju Ma, Donald Azuatalam, Thomas Power et al.
Battery storage, particularly residential battery storage coupled with rooftop PV, is emerging as an essential component of the smart grid technology mix. However, including battery storage and other flexible resources like electric vehicles and loads with thermal inertia into a probabilistic analysis based on Monte Carlo (MC) simulation is challenging, because their operational profiles are determined by computationally intensive optimization. Additionally, MC analysis requires a large pool of statistically-representative demand profiles to sample from. As a result, the analysis of the network impact of PV-battery systems has attracted little attention in the existing literature. To fill these knowledge gaps, this paper proposes a novel probabilistic framework to study the impact of PV-battery systems on low-voltage distribution networks. Specifically, the framework incorporates home energy management(HEM) operational decisions within the MC time series power flow analysis. First, using available smart meter data, we use a Bayesian nonparametric model to generate statistically-representative synthetic demand and PV profiles. Second, a policy function approximation that emulates battery scheduling decisions is used to make the simulation of optimization-based HEM feasible within the MC framework. The efficacy of our method is demonstrated on three representative low-voltage feeders, where the computation time to execute our MC framework is 5% of that when using explicit optimization methods in each MC sample. The assessment results show that uncoordinated battery scheduling has a limited beneficial impact, which is against the conjecture that batteries will serendipitously mitigate the technical problems induced by PV generation.
1.2SYSep 20, 2018
Impacts of Community and Distributed Energy Storage Systems on Unbalanced Low Voltage NetworksYiju Ma, Mohammad Seydali Seyf Abad, Donald Azuatalam et al.
Energy storage systems (EES) are expected to be an indispensable resource for mitigating the effects on networks of high penetrations of distributed generation in the near future. This paper analyzes the benefits of EES in unbalanced low voltage (LV) networks regarding three aspects, namely, power losses, the hosting capacity and network unbalance. For doing so, a mixed integer quadratic programmming model (MIQP) is developed to minimize annual energy losses and determine the sizing and placement of ESS, while satisfying voltage constraints. A real unbalanced LV UK grid is adopted to examine the effects of ESS under two scenarios: the installation of one community ESS (CESS) and multiple distributed ESSs (DESSs). The results illustrate that both scenarios present high performance in accomplishing the above tasks, while DESSs, with the same aggregated size, are slightly better. This margin is expected to be amplified as the aggregated size of DESSs increases.
2.3SYSep 19, 2018
Decentralized P2P Energy Trading under Network Constraints in a Low-Voltage NetworkJaysson Guerrero, Archie Chapman, Gregor Verbic
The increasing uptake of distributed energy resources (DERs) in distribution systems and the rapid advance of technology have established new scenarios in the operation of low-voltage networks. In particular, recent trends in cryptocurrencies and blockchain have led to a proliferation of peer-to-peer (P2P) energy trading schemes, which allow the exchange of energy between the neighbors without any intervention of a conventional intermediary in the transactions. Nevertheless, far too little attention has been paid to the technical constraints of the network under this scenario. A major challenge to implementing P2P energy trading is that of ensuring that network constraints are not violated during the energy exchange. This paper proposes a methodology based on sensitivity analysis to assess the impact of P2P transactions on the network and to guarantee an exchange of energy that does not violate network constraints. The proposed method is tested on a typical UK low-voltage network. The results show that our method ensures that energy is exchanged between users under the P2P scheme without violating the network constraints, and that users can still capture the economic benefits of the P2P architecture.
1.2SYAug 1, 2017
A Framework for Frequency Stability Assessment of Future Power Systems: An Australian Case StudyAhmad Shabir Ahmadyar, Shariq Riaz, Gregor Verbic et al.
The increasing penetration of non-synchronous renewable energy sources (NS-RES) alters the dynamic characteristic, and consequently, the frequency behaviour of a power system. To accurately identify these changing trends and address them in a systematic way, it is necessary to assess a large number of scenarios. Given this, we propose a frequency stability assessment framework based on a time-series approach that facilitates the analysis of a large number of future power system scenarios. We use this framework to assess the frequency stability of the Australian future power system by considering a large number of future scenarios and sensitivity of different parameters. By doing this, we identify a maximum non-synchronous instantaneous penetration range from the frequency stability point of view. Further, to reduce the detrimental impacts of high NS-RES penetration on system frequency stability, a dynamic inertia constraint is derived and incorporated in the market dispatch model. The results show that such a constraint guarantees frequency stability of the system for all credible contingencies. Also, we assess and quantify the contribution of synchronous condensers, synthetic inertia of wind farms and a governor-like response from de-loaded wind farms on system frequency stability. The results show that the last option is the most effective one.
24.9AIApr 9, 2012
Knapsack based Optimal Policies for Budget-Limited Multi-Armed BanditsLong Tran-Thanh, Archie Chapman, Alex Rogers et al.
In budget-limited multi-armed bandit (MAB) problems, the learner's actions are costly and constrained by a fixed budget. Consequently, an optimal exploitation policy may not be to pull the optimal arm repeatedly, as is the case in other variants of MAB, but rather to pull the sequence of different arms that maximises the agent's total reward within the budget. This difference from existing MABs means that new approaches to maximising the total reward are required. Given this, we develop two pulling policies, namely: (i) KUBE; and (ii) fractional KUBE. Whereas the former provides better performance up to 40% in our experimental settings, the latter is computationally less expensive. We also prove logarithmic upper bounds for the regret of both policies, and show that these bounds are asymptotically optimal (i.e. they only differ from the best possible regret by a constant factor).
2.3GTMar 15, 2012
Automated Planning in Repeated Adversarial GamesEnrique Munoz de Cote, Archie C. Chapman, Adam M. Sykulski et al.
Game theory's prescriptive power typically relies on full rationality and/or self-play interactions. In contrast, this work sets aside these fundamental premises and focuses instead on heterogeneous autonomous interactions between two or more agents. Specifically, we introduce a new and concise representation for repeated adversarial (constant-sum) games that highlight the necessary features that enable an automated planing agent to reason about how to score above the game's Nash equilibrium, when facing heterogeneous adversaries. To this end, we present TeamUP, a model-based RL algorithm designed for learning and planning such an abstraction. In essence, it is somewhat similar to R-max with a cleverly engineered reward shaping that treats exploration as an adversarial optimization problem. In practice, it attempts to find an ally with which to tacitly collude (in more than two-player games) and then collaborates on a joint plan of actions that can consistently score a high utility in adversarial repeated games. We use the inaugural Lemonade Stand Game Tournament to demonstrate the effectiveness of our approach, and find that TeamUP is the best performing agent, demoting the Tournament's actual winning strategy into second place. In our experimental analysis, we show hat our strategy successfully and consistently builds collaborations with many different heterogeneous (and sometimes very sophisticated) adversaries.
1.2GTFeb 14, 2012
Filtered Fictitious Play for Perturbed Observation Potential Games and Decentralised POMDPsArchie C. Chapman, Simon A. Williamson, Nicholas R. Jennings
Potential games and decentralised partially observable MDPs (Dec-POMDPs) are two commonly used models of multi-agent interaction, for static optimisation and sequential decisionmaking settings, respectively. In this paper we introduce filtered fictitious play for solving repeated potential games in which each player's observations of others' actions are perturbed by random noise, and use this algorithm to construct an online learning method for solving Dec-POMDPs. Specifically, we prove that noise in observations prevents standard fictitious play from converging to Nash equilibrium in potential games, which also makes fictitious play impractical for solving Dec-POMDPs. To combat this, we derive filtered fictitious play, and provide conditions under which it converges to a Nash equilibrium in potential games with noisy observations. We then use filtered fictitious play to construct a solver for Dec-POMDPs, and demonstrate our new algorithm's performance in a box pushing problem. Our results show that we consistently outperform the state-of-the-art Dec-POMDP solver by an average of 100% across the range of noise in the observation function.