Data Duplication Detection by Applying Similarity Strategies |
Author(s): |
| B. Lalitha , Sri GVG Visalakshi college for women; S. Shobana, Sri GVG Visalakshi college for women |
Keywords: |
| De-duplication; Similarity Strategies; Sorted Neighborhood Method (SNM); Windowing; Blocking |
Abstract |
|
Record deduplication is the task of identifying a data repository or records that refer to the same real world entity or object in spite of misspelling words, types, different schema representations or data types. The data gathered from numerous resources might have quality problems in it. The concept to identify duplicates by using windowing and blocking strategy is to achieve better precision, good efficiency and also to reduce the false positive rate all are in accordance with the estimated similarities of records. Various Similarity metrics area unit unremarkably accustomed acknowledge the similar field entries. So the main focus of this project is to apply appropriate similarity measure on appropriate data to properly identifying the duplicates. De-duplication is a property which provides additional information of similarities between the two entities. Thus, in today’s knowledge centrical atmosphere there live Brobdingnagian numbers of defects in similarity measure. As a result to identify the duplicates is always been a challenging task. In this project the primary focus is given on exact identification of duplicates in the database by applying concept of windowing & blocking. The objective is to achieve better precision, good efficiency and also to reduce the false positive rate all are in accordance with the estimated similarity strategies of records. |
Other Details |
|
Paper ID: IJSRDV7I10221 Published in: Volume : 7, Issue : 1 Publication Date: 01/04/2019 Page(s): 252-254 |
Article Preview |
|
|
|
|
