1.2SYNov 20, 2015
On Cooperative Behavior of Open Homogeneous Chemical Reaction Systems in the Extent DomainNirav Bhatt, Sriniketh Srinivasan
Material balance equations describe the dynamics of the species in open reaction systems and contain information regarding reaction topology, kinetics and operation mode. For reaction systems, the state variables (the numbers of moles, or concentrations) have recently been transformed into decoupled reaction variants (extents of reaction), and reaction invariants (extents of flow) (Amrhein et al., AIChE Journal, 2010). This paper analyses the conditions under which an open homogeneous reaction system is cooperative in the extents domain. Further, it is shown that the dynamics of the extents of flow exhibit cooperative behavior. Further, we provide the conditions under which the dynamics of the extents of reaction exhibit cooperative behavior. Our results provide physical insights into cooperative and competitive nature of the underlying reaction system in the presence of material exchange with surrounding (i.e., inlet and outlet flows). The results of the article are demonstrated via examples.
1.4LGFeb 11
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman Ravindran
This work addresses the problem of offline safe imitation learning (IL), where the goal is to learn safe and reward-maximizing policies from demonstrations that do not have per-timestep safety cost or reward information. In many real-world domains, online learning in the environment can be risky, and specifying accurate safety costs can be difficult. However, it is often feasible to collect trajectories that reflect undesirable or unsafe behavior, implicitly conveying what the agent should avoid. We refer to these as non-preferred trajectories. We propose a novel offline safe IL algorithm, OSIL, that infers safety from non-preferred demonstrations. We formulate safe policy learning as a Constrained Markov Decision Process (CMDP). Instead of relying on explicit safety cost and reward annotations, OSIL reformulates the CMDP problem by deriving a lower bound on reward maximizing objective and learning a cost model that estimates the likelihood of non-preferred behavior. Our approach allows agents to learn safe and reward-maximizing behavior entirely from offline demonstrations. We empirically demonstrate that our approach can learn safer policies that satisfy cost constraints without degrading the reward performance, thus outperforming several baselines.
4.1LGNov 11, 2025
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman Ravindran
In this work, we study the problem of offline safe imitation learning (IL). In many real-world settings, online interactions can be risky, and accurately specifying the reward and the safety cost information at each timestep can be difficult. However, it is often feasible to collect trajectories reflecting undesirable or risky behavior, implicitly conveying the behavior the agent should avoid. We refer to these trajectories as non-preferred trajectories. Unlike standard IL, which aims to mimic demonstrations, our agent must also learn to avoid risky behavior using non-preferred trajectories. In this paper, we propose a novel approach, SafeMIL, to learn a parameterized cost that predicts if the state-action pair is risky via Multiple Instance Learning. The learned cost is then used to avoid non-preferred behaviors, resulting in a policy that prioritizes safety. We empirically demonstrate that our approach can learn a safer policy that satisfies cost constraints without degrading the reward performance, thereby outperforming several baselines.
5.0ROMay 30, 2023
GAN-MPC: Training Model Predictive Controllers with Parameterized Cost Functions using Demonstrations from Non-identical ExpertsReturaj Burnwal, Anirban Santara, Nirav P. Bhatt et al.
Model predictive control (MPC) is a popular approach for trajectory optimization in practical robotics applications. MPC policies can optimize trajectory parameters under kinodynamic and safety constraints and provide guarantees on safety, optimality, generalizability, interpretability, and explainability. However, some behaviors are complex and it is difficult to hand-craft an MPC objective function. A special class of MPC policies called Learnable-MPC addresses this difficulty using imitation learning from expert demonstrations. However, they require the demonstrator and the imitator agents to be identical which is hard to satisfy in many real world applications of robotics. In this paper, we address the practical problem of training Learnable-MPC policies when the demonstrator and the imitator do not share the same dynamics and their state spaces may have a partial overlap. We propose a novel approach that uses a generative adversarial network (GAN) to minimize the Jensen-Shannon divergence between the state-trajectory distributions of the demonstrator and the imitator. We evaluate our approach on a variety of simulated robotics tasks of DeepMind Control suite and demonstrate the efficacy of our approach at learning the demonstrator's behavior without having to copy their actions.
1.8LGMay 21, 2019
Learning Conserved Networks from FlowsSatya Jayadev P., Shankar Narasimhan, Nirav Bhatt
A challenging problem in complex networks is the network reconstruction problem from data. This work deals with a class of networks denoted as conserved networks, in which a flow associated with every edge and the flows are conserved at all non-source and non-sink nodes. We propose a novel polynomial time algorithm to reconstruct conserved networks from flow data by exploiting graph theoretic properties of conserved networks combined with learning techniques. We prove that exact network reconstruction is possible for arborescence networks. We also extend the methodology for reconstructing networks from noisy data and explore the reconstruction performance on arborescence networks with different structural characteristics.
2.1LGNov 19, 2015
A Novel Approach for Phase Identification in Smart Grids Using Graph Theory and Principal Component AnalysisP Satya Jayadev, Aravind Rajeswaran, Nirav P Bhatt et al.
Consumers with low demand, like households, are generally supplied single-phase power by connecting their service mains to one of the phases of a distribution transformer. The distribution companies face the problem of keeping a record of consumer connectivity to a phase due to uninformed changes that happen. The exact phase connectivity information is important for the efficient operation and control of distribution system. We propose a new data driven approach to the problem based on Principal Component Analysis (PCA) and its Graph Theoretic interpretations, using energy measurements in equally timed short intervals, generated from smart meters. We propose an algorithm for inferring phase connectivity from noisy measurements. The algorithm is demonstrated using simulated data for phase connectivities in distribution networks.