NI LGJan 10, 2021

Learning Augmented Index Policy for Optimal Service Placement at the Network Edge

arXiv:2101.03641v28.010 citations

Originality Incremental advance

AI Analysis

This work provides adaptive algorithms for network operators to optimize service placement at the edge, aiming to reduce customer latency, which is an incremental improvement to existing methods.

This paper addresses the problem of optimally placing services at the network edge to minimize average service delivery latency for customers. The authors formulate this as a Markov Decision Process (MDP) and, to overcome the curse of dimensionality, derive Whittle indices for a single-service MDP, which exhibits a threshold structure. They then develop two learning-augmented algorithms, UCB-Whittle and Q-learning-Whittle, to handle unknown and time-varying rates, showing excellent empirical performance in simulations.

We consider the problem of service placement at the network edge, in which a decision maker has to choose between $N$ services to host at the edge to satisfy the demands of customers. Our goal is to design adaptive algorithms to minimize the average service delivery latency for customers. We pose the problem as a Markov decision process (MDP) in which the system state is given by describing, for each service, the number of customers that are currently waiting at the edge to obtain the service. However, solving this $N$-services MDP is computationally expensive due to the curse of dimensionality. To overcome this challenge, we show that the optimal policy for a single-service MDP has an appealing threshold structure, and derive explicitly the Whittle indices for each service as a function of the number of requests from customers based on the theory of Whittle index policy. Since request arrival and service delivery rates are usually unknown and possibly time-varying, we then develop efficient learning augmented algorithms that fully utilize the structure of optimal policies with a low learning regret. The first of these is UCB-Whittle, and relies upon the principle of optimism in the face of uncertainty. The second algorithm, Q-learning-Whittle, utilizes Q-learning iterations for each service by using a two time scale stochastic approximation. We characterize the non-asymptotic performance of UCB-Whittle by analyzing its learning regret, and also analyze the convergence properties of Q-learning-Whittle. Simulation results show that the proposed policies yield excellent empirical performance.

View on arXiv PDF

Similar