B. Lantz

3.3NIApr 29, 2025

Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning

Jinsun Yoo, ChonLam Lao, Lianjie Cao et al.

This paper lays the foundation for Genie, a testing framework that captures the impact of real hardware network behavior on ML workload performance, without requiring expensive GPUs. Genie uses CPU-initiated traffic over a hardware testbed to emulate GPU to GPU communication, and adapts the ASTRA-sim simulator to model interaction between the network and the ML workload.

B. Lantz

1 Paper