Clean Me If You Can: A Large Collection of Real-World Addresses for Data Cleaning Benchmarking
For data cleaning researchers, this dataset exposes the limitations of current methods on realistic dirty data.
The authors provide a large real-world dataset of postal addresses with ground truth for data cleaning benchmarking, and show that existing cleaning approaches perform poorly on it.
There has been extensive research on automating and scaling data cleaning, i.e., the detection and correction of erroneous values in tabular data. Yet, existing approaches often perform well only within controlled environments. One of the major bottlenecks in data cleaning research is the lack of real-world datasets. In this paper, we address this gap by providing a large, dirty dataset with postal entries and their corresponding ground truth. We discuss the design decisions and challenges for obtaining the dataset. We demonstrate the limitations of existing cleaning approaches when faced with our proposed datasets and derive guidelines for future research.