Scalable Matching and Clustering of Entities with FAMER

Alieh Saeedi, Markus Nentwig, Eric Peukert, Erhard Rahm

Abstract


Entity resolution identifies semantically equivalent entities, e.g. describing the same product or customer. It is especially challenging for Big Data applications where large volumes of data from many sources have to be matched and integrated. We therefore introduce a scalable entity resolution framework called FAMER (FAst Multi-source Entity Resolution system) that is based on Apache Flink for distributed execution and that can holistically match entities from multiple sources. For the latter purpose, FAMER includes multiple clustering schemes that group matching entities from different sources within clusters. In addition to previously known clustering schemes FAMER includes new approaches tailored to multi-source entity resolution. We perform a detailed comparative evaluation of eight clustering schemes for different real-life and synthetically generated datasets. The evaluation considers both the match quality as well as the scalability for different numbers of machines and data sizes.


Keywords:

Clustering; Matching; Distributed processing; Multi-source

Full Text:

PDF


DOI: 10.7250/csimq.2018-16.04

Cited-By

1. Big Data Competence Center ScaDS Dresden/Leipzig: Overview and selected research activities
Erhard Rahm, Wolfgang E. Nagel, Eric Peukert, René Jäkel, Fabian Gärtner, Peter F. Stadler, Daniel Wiegreffe, Dirk Zeckzer, Wolfgang Lehner
Datenbank-Spektrum  vol: 19  issue: 1  first page: 5  year: 2019  
doi: 10.1007/s13222-018-00303-6

Refbacks

  • There are currently no refbacks.