Scalable Matching and Clustering of Entities with FAMER
Abstract
Entity resolution identifies semantically equivalent entities, e.g. describing the same product or customer. It is especially challenging for Big Data applications where large volumes of data from many sources have to be matched and integrated. We therefore introduce a scalable entity resolution framework called FAMER (FAst Multi-source Entity Resolution system) that is based on Apache Flink for distributed execution and that can holistically match entities from multiple sources. For the latter purpose, FAMER includes multiple clustering schemes that group matching entities from different sources within clusters. In addition to previously known clustering schemes FAMER includes new approaches tailored to multi-source entity resolution. We perform a detailed comparative evaluation of eight clustering schemes for different real-life and synthetically generated datasets. The evaluation considers both the match quality as well as the scalability for different numbers of machines and data sizes.
Keywords: |
Clustering; Matching; Distributed processing; Multi-source
|
Full Text: |
DOI: 10.7250/csimq.2018-16.04
Cited-By
1. EEUPL: Towards effective and efficient user profile linkage across multiple social platforms
Manman Wang, Weiqing Wang, Wei Chen, Lei Zhao
World Wide Web vol: 24 issue: 5 first page: 1731 year: 2021
doi: 10.1007/s1128
Refbacks
- There are currently no refbacks.
Copyright (c) 2018 Complex Systems Informatics and Modeling Quarterly