Artificial Intelligence (AI)-Centric Management of Resources in Modern Distributed Computing Systems
This work tackles resource management inefficiencies in distributed systems like cloud data centers, but it is incremental as it primarily reviews and conceptualizes existing AI approaches rather than introducing new methods.
The paper addresses the challenge of managing resources in modern distributed computing systems, which are complex and dynamic, by advocating for AI-centric, data-driven solutions, and demonstrates feasibility through real-time use cases from Google Cloud and Microsoft Azure.
Contemporary Distributed Computing Systems (DCS) such as Cloud Data Centres are large scale, complex, heterogeneous, and distributed across multiple networks and geographical boundaries. On the other hand, the Internet of Things (IoT)-driven applications are producing a huge amount of data that requires real-time processing and fast response. Managing these resources efficiently to provide reliable services to end-users or applications is a challenging task. The existing Resource Management Systems (RMS) rely on either static or heuristic solutions inadequate for such composite and dynamic systems. The advent of Artificial Intelligence (AI) due to data availability and processing capabilities manifested into possibilities of exploring data-driven solutions in RMS tasks that are adaptive, accurate, and efficient. In this regard, this paper aims to draw the motivations and necessities for data-driven solutions in resource management. It identifies the challenges associated with it and outlines the potential future research directions detailing where and how to apply the data-driven techniques in the different RMS tasks. Finally, it provides a conceptual data-driven RMS model for DCS and presents the two real-time use cases (GPU frequency scaling and data centre resource management from Google Cloud and Microsoft Azure) demonstrating AI-centric approaches' feasibility.