2.3NAOct 28, 2011
The Dominant Eigenvalue of an Essentially Nonnegative TensorLiping Zhang, Liqun Qi, Ziyan Luo
It is well known that the dominant eigenvalue of a real essentially nonnegative matrix is a convex function of its diagonal entries. This convexity is of practical importance in population biology, graph theory, demography, analytic hierarchy process and so on. In this paper, the concept of essentially nonnegativity is extended from matrices to higher order tensors, and the convexity and log convexity of dominant eigenvalues for such a class of tensors are established. Particularly, for any nonnegative tensor, the spectral radius turns out to be the dominant eigenvalue and hence possesses these convexities. Finally, an algorithm is given to calculate the dominant eigenvalue, and numerical results are reported to show the effectiveness of the proposed algorithm.
2.3NAFeb 29, 2012
M-tensors and The Positive Definiteness of a Multivariate FormLiping Zhang, Liqun Qi, Guanglu Zhou
We study M-tensors and various properties of M-tensors are given. Specially, we show that the smallest real eigenvalue of M-tensor is positive corresponding to a nonnegative eigenvector. We propose an algorithm to find the smallest positive eigenvalue and then apply the property to study the positive definiteness of a multivariate form.
4.1LGNov 6, 2025
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement LearningJiaming Zhang, Yujie Yang, Haoning Wang et al.
Safe reinforcement learning (safe RL) aims to respect safety requirements while optimizing long-term performance. In many practical applications, however, the problem involves an infinite number of constraints, known as semi-infinite safe RL (SI-safe RL). Such constraints typically appear when safety conditions must be enforced across an entire continuous parameter space, such as ensuring adequate resource distribution at every spatial location. In this paper, we propose exchange policy optimization (EPO), an algorithmic framework that achieves optimal policy performance and deterministic bounded safety. EPO works by iteratively solving safe RL subproblems with finite constraint sets and adaptively adjusting the active set through constraint expansion and deletion. At each iteration, constraints with violations exceeding the predefined tolerance are added to refine the policy, while those with zero Lagrange multipliers are removed after the policy update. This exchange rule prevents uncontrolled growth of the working set and supports effective policy training. Our theoretical analysis demonstrates that, under mild assumptions, strategies trained via EPO achieve performance comparable to optimal solutions with global constraint violations strictly remaining within a prescribed bound.