High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Enhancing the Scalability and Efficiency of Big Data using Clustering Algorithms

Author(s):

Kosha Kothari , L.J Institute of Technology,Ahmedabad; Ompriya Kale, L.J Institute of Technology,Ahmedabad

Keywords:

Clustering Algorithms, Unsupervised learning, Data Mining, K-Means, DBSCAN, Big data

Abstract

The data that is been produced by numerous scientific applications and incorporated environment has grown rapidly not only in size but also in variety in current era. The data collected is of very large amount and there is an adversity in gathering and evaluating such big data. The main goal of clustering is to categorize data into clusters such that objects are grouped in the same cluster when they are “similar” according to similarities, traits and behavior. The effectiveness and efficiency of the existing algorithms is, somewhat limited, since clustering with big data requires clustering high-dimensional feature vectors and since big data often contain large amounts of noise and large datasets. K-Means and DBSCAN are complement to analyze big data on cloud environment. A hybrid approach based on parallel K-Means and parallel DBSCAN is proposed to overcome the drawbacks of both these algorithms. The hybrid approach combines the benefits of both the clustering techniques. The proposed technique is evaluated on the MapReduce framework of Hadoop Platform. The results show that the proposed hybrid approach is an improved version of parallel K-Means clustering algorithm and parallel DBSCAN algorithm.

Other Details

Paper ID: IJSRDV3I30504
Published in: Volume : 3, Issue : 3
Publication Date: 01/06/2015
Page(s): 3594-3598

Article Preview

Download Article