NIApr 29, 2025
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine LearningJinsun Yoo, ChonLam Lao, Lianjie Cao et al.
This paper lays the foundation for Genie, a testing framework that captures the impact of real hardware network behavior on ML workload performance, without requiring expensive GPUs. Genie uses CPU-initiated traffic over a hardware testbed to emulate GPU to GPU communication, and adapts the ASTRA-sim simulator to model interaction between the network and the ML workload.