DBJul 1

From Single to Multiple Attributes: Experimental Insights on Sampling-Based Distinct Combination Estimation in GROUP-BY Queries

arXiv:2607.008681.9
Predicted impact top 96% in DB · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work provides empirical insights for database practitioners and researchers on the limitations of current sampling-based methods for multi-attribute GROUP-BY cardinality estimation.

The paper addresses the underexplored challenge of estimating distinct combinations in multi-attribute GROUP-BY queries. Through empirical evaluation on real-world datasets and TPC-H, they find that existing sampling-based methods often fail to provide accurate estimates, especially for filtered queries, and offer recommendations for future estimator design.

Estimating the number of distinct combinations in multi-attribute GROUP-BY queries remains a significant yet underexplored challenge. Current cardinality estimation techniques primarily focus on SPJ queries (i.e., selections, projections, and joins) and neglect GROUP-BY operations; meanwhile, distinct value estimation research has mainly targeted the single-attribute setting. Although sampling-based methods, including recent approaches with learned models, can theoretically support multi-attribute estimation, their practical effectiveness remains unclear. A comprehensive empirical evaluation is thus lacking to address whether joint distribution information from samples alone is sufficient for accurate multi-attribute estimation, whether existing methods fully exploit single-attribute information and can be further optimized, and whether filtered GROUP-BY queries can be accurately estimated. To this end, we propose a specialized workload generator for multi-attribute GROUP-BY queries and generate both filtered and non-filtered queries over four real-world datasets. By evaluating existing methods across synthetic workloads and the multi-table TPC-H benchmark, we analyze the sources of GROUP-BY cardinality estimation errors and their impact on PostgreSQL's plan selection, offering key recommendations for future estimator design.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes