Scalable Matching and Clustering of Entities with FAMER

Authors

DOI:

https://doi.org/10.7250/csimq.2018-16.04

Keywords:

Clustering, Matching, Distributed processing, Multi-source

Abstract

Entity resolution identifies semantically equivalent entities, e.g. describing the same product or customer. It is especially challenging for Big Data applications where large volumes of data from many sources have to be matched and integrated. We therefore introduce a scalable entity resolution framework called FAMER (FAst Multi-source Entity Resolution system) that is based on Apache Flink for distributed execution and that can holistically match entities from multiple sources. For the latter purpose, FAMER includes multiple clustering schemes that group matching entities from different sources within clusters. In addition to previously known clustering schemes FAMER includes new approaches tailored to multi-source entity resolution. We perform a detailed comparative evaluation of eight clustering schemes for different real-life and synthetically generated datasets. The evaluation considers both the match quality as well as the scalability for different numbers of machines and data sizes.

Downloads

Published

31.10.2018

How to Cite

Saeedi, A., Nentwig, M., Peukert, E., & Rahm, E. (2018). Scalable Matching and Clustering of Entities with FAMER. Complex Systems Informatics and Modeling Quarterly, 16, 61-83. https://doi.org/10.7250/csimq.2018-16.04