DBJun 30

Clean Me If You Can: A Large Collection of Real-World Addresses for Data Cleaning Benchmarking

arXiv:2606.319832.6
Predicted impact top 91% in DB · last 90 daysOriginality Synthesis-oriented
AI Analysis

For data cleaning researchers, this dataset exposes the limitations of current methods on realistic dirty data.

The authors provide a large real-world dataset of postal addresses with ground truth for data cleaning benchmarking, and show that existing cleaning approaches perform poorly on it.

There has been extensive research on automating and scaling data cleaning, i.e., the detection and correction of erroneous values in tabular data. Yet, existing approaches often perform well only within controlled environments. One of the major bottlenecks in data cleaning research is the lack of real-world datasets. In this paper, we address this gap by providing a large, dirty dataset with postal entries and their corresponding ground truth. We discuss the design decisions and challenges for obtaining the dataset. We demonstrate the limitations of existing cleaning approaches when faced with our proposed datasets and derive guidelines for future research.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes