High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Analysis on Smart Removal of Redundant Data Using Progressive Techniques

Author(s):

K. D. Mane , S.B. Patil C.O.E. Indapur; M. H. More, S.B. Patil C.O.E. Indapur; V. B. Bansode, S.B. Patil C.O.E. Indapur; P. S. Gavade, S.B. Patil C.O.E. Indapur

Keywords:

Data Cleaning, Data Duplication, Progressiveness

Abstract

Data are among the most important assets of a company. but due to data changes and sloppy data entry, errors such as duplicate entries might occur, making data clean sing and in particular duplicate detection indispensable. However, the poor size of today’s data sets renders duplicate detection processes expensive. Online retailers, fore.g., offer hug catalogs comprising a constantly growing set of items from many different suppliers. As independent persons change the product portfolio, duplicate rise. Although there is an obvious need for de-duplication. Progressive duplicate detection identifies most duplicate pairs early in the detection process. Instead of reducing the overall time needed to finish the entire process, progressive approaches try to reduce the average time after which a duplicate is found .early termination, in particular, then yields more complete results on a progressive algorithm than on any traditional approach.

Other Details

Paper ID: IJSRDV5I20823
Published in: Volume : 5, Issue : 2
Publication Date: 01/05/2017
Page(s): 1973-1975

Article Preview

Download Article