ARJun 10

Making Locality-aware GEMM Compatible with Page-Granularity Placement on Chiplet GPUs

arXiv:2606.11718v18.0h-index: 3
Predicted impact top 36% in AR · last 90 daysOriginality Incremental advance
AI Analysis

For GPU architects and system software designers, this work addresses the incompatibility between locality-aware scheduling and page-granularity memory placement in multi-chiplet GPUs, improving performance and energy efficiency for LLM workloads.

Chiplet-Contiguous Layout enables locality-aware data placement for GEMM on chiplet GPUs, reducing remote HBM traffic by 24.7x on Qwen 3 30B and 19.2x on Llama 3.1 70B over 4KB interleaving.

Multi-chiplet GPUs scale compute throughput and high-bandwidth memory (HBM) capacity, but their non-uniform memory system makes locality between chiplets and their data critical to the GPU's performance and energy efficiency. Locality-aware scheduling and data placement identify which data should reside near each chiplet. However, in general matrix multiplication (GEMM), locality-aware data placement often becomes incompatible with a fixed page-granularity data interleaving, since the optimal granularity for mapping data across chiplets varies widely across workloads. We propose Chiplet-Contiguous Layout, a global memory layout that stores chiplet-local data contiguously. Chiplet-Contiguous Layout enables locality-aware placement compatible with page-granularity placement across diverse large language model (LLM) GEMM shapes, without changes to the operating system or hardware. On representative LLM inference and training GEMMs from Qwen 3 30B and Llama 3.1 70B, Chiplet-Contiguous Layout on average reduces remote HBM traffic by 24.7x on Qwen and 19.2x on Llama over 4KB interleaving, and by 4.1x and 2.1x over coarse locality-aware placement.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes