Analysis on Smart Removal of Redundant Data Using Progressive Techniques |
Author(s): |
| K. D. Mane , S.B. Patil C.O.E. Indapur; M. H. More, S.B. Patil C.O.E. Indapur; V. B. Bansode, S.B. Patil C.O.E. Indapur; P. S. Gavade, S.B. Patil C.O.E. Indapur |
Keywords: |
| Data Cleaning, Data Duplication, Progressiveness |
Abstract |
|
Data are among the most important assets of a company. but due to data changes and sloppy data entry, errors such as duplicate entries might occur, making data clean sing and in particular duplicate detection indispensable. However, the poor size of today’s data sets renders duplicate detection processes expensive. Online retailers, fore.g., offer hug catalogs comprising a constantly growing set of items from many different suppliers. As independent persons change the product portfolio, duplicate rise. Although there is an obvious need for de-duplication. Progressive duplicate detection identiï¬es most duplicate pairs early in the detection process. Instead of reducing the overall time needed to ï¬nish the entire process, progressive approaches try to reduce the average time after which a duplicate is found .early termination, in particular, then yields more complete results on a progressive algorithm than on any traditional approach. |
Other Details |
|
Paper ID: IJSRDV5I20823 Published in: Volume : 5, Issue : 2 Publication Date: 01/05/2017 Page(s): 1973-1975 |
Article Preview |
|
|
|
|
