Zhonggen Li

h-index2
3papers
8citations

3 Papers

7.6DCMar 12Code
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration

Zhonggen Li, Xiangyu Ke, Yifan Zhu et al.

Graph embeddings map graph nodes to continuous vectors and are foundational to community detection, recommendation, and many scientific applications. At billion-scale, however, existing graph embedding systems face a trade-off: they either rely on large in-memory footprints across many GPUs (limited scalability) or repeatedly stream data from disk (incurring severe I/O overhead and low GPU utilization). In this paper, we propose Legend, a lightweight heterogeneous system for graph embedding that systematically redesigns data management across CPU, GPU, and NVMe SSD resources. Legend combines three practical ideas: (1) a prefetch-friendly embedding-loading order that lets GPUs efficiently prefetch necessary embeddings directly from NVMe SSD with low I/O amplification; (2) a high-throughput GPU-SSD direct-access driver tuned for the access patterns of embedding training; and (3) a customized parallel execution strategy that maximizes GPU utilization. Together, these components let Legend store and stream vast embedding data without overprovisioning GPU memory or suffering I/O stalls. Extensive experiments on billion-scale graphs demonstrate that Legend speeds up end-to-end workloads by up to 4.8x versus state-of-the-art systems, and matches their performance on the largest workloads while using only one quarter of the GPUs.

6.1DBJun 15
Accelerating High-Dimensional Nearest Neighbor Search with Dynamic Query Preference

Yifan Zhu, Ruijie Zhao, Zhonggen Li et al.

Approximate Nearest Neighbor Search (ANNS) has emerged as an essential operation in modern database and AI systems. While graph-based methods like NSG demonstrate state-of-the-art ANNS performance, they typically ignore that query distributions are often skewed. In real-world scenarios, user preferences and time-varying access patterns lead to non-uniform workloads, where specific data regions are retrieved significantly more frequently than others. Meanwhile, these patterns evolve over time, making pre-built indexes outdated and thus inefficient for future query workloads. Motivated by this, we propose DQF, a novel Dual-Index Query Framework for dynamic query preference. This dual-index structure comprises a Hot Index containing frequently accessed nodes and a Full Index covering the entire dataset, so that hot queries can be answered faster within the compact Hot Index while cold queries still obtain complete results from the Full Index. Furthermore, we propose a three-phase competitive search in which both layers share a single priority queue. A lightweight decision tree detects when the top-k results have stabilized and triggers per-query early termination. To address temporal shifts in query patterns, we design an adaptive update mechanism that periodically promotes new high-frequency nodes to the Hot Index while demoting outdated ones. Experiments on five real-world datasets demonstrate that DQF achieves a 2.2-6.9x speedup over the strongest baseline on each million-scale dataset at 95% recall. Moreover, it scales to 100M vectors with consistent performance gains, successfully adapting to distribution shifts without requiring Full Index reconstruction.

8.9DBApr 22
A GPU-Accelerated Framework for Multi-Attribute Range Filtered Approximate Nearest Neighbor Search

Zhonggen Li, Haoran Yu, Yifan Zhu et al.

Range-filtered approximate nearest neighbor search (RFANNS) is increasingly critical for modern vector databases. However, existing solutions suffer from severe index inflation and construction overhead. Furthermore, they rely exclusively on CPUs for the heavy indexing and query processing, failing to leverage the powerful computational capabilities of GPUs. In this paper, we present Garfield, a GPU-accelerated framework for multi-attribute range filtered ANNS that overcomes these bottlenecks through designing a lightweight index structure and hardware-aware execution pipeline. Garfield introduces the GMG index, which partitions data into cells and builds local graph indexes. By adding a constant number of cross-cell edges, it guarantees linear storage and indexing overhead. For queries, Garfield utilizes a cluster-guided ordering strategy that reorders query-relevant cells, enabling a highly efficient cell-by-cell traversal on the GPU that aggressively reuses candidates as entry points across cells. To handle datasets exceeding GPU memory, Garfield features a cell-oriented out-of-core pipeline. It dynamically schedules cells to minimize the number of active queries per batch and overlaps GPU computation with CPU-to-GPU index streaming. Extensive evaluations demonstrate that Garfield reduces index size by 4.4x, while delivering 119.8x higher throughput than state-of-the-art RFANNS methods.