Back to Explore
cs.DCComputer Science

Distributed Computing

Distributed systems, parallel computing, cloud

20.5AIMay 7Code87
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?

Keisuke Kamahori, Shihang Li, Simon Peter et al.

This work challenges the paradigm of general-purpose LLM serving stacks by proposing generation-time specialization, which could benefit system builders and researchers dealing with diverse model architectures, workloads, and hardware.

27.7DCMay 11
Accelerating Compound LLM Training Workloads with Maestro

Xiulong Yuan, Hongqing Chen, Jiaxuan Peng et al.

For practitioners training complex multi-component LLMs, Maestro provides a practical framework that significantly improves GPU utilization and throughput over monolithic approaches.