Enhancing the Scalability and Efficiency of Big Data using Clustering Algorithms |
Author(s): |
| Kosha Kothari , L.J Institute of Technology,Ahmedabad; Ompriya Kale, L.J Institute of Technology,Ahmedabad |
Keywords: |
| Clustering Algorithms, Unsupervised learning, Data Mining, K-Means, DBSCAN, Big data |
Abstract |
|
The data that is been produced by numerous scientific applications and incorporated environment has grown rapidly not only in size but also in variety in current era. The data collected is of very large amount and there is an adversity in gathering and evaluating such big data. The main goal of clustering is to categorize data into clusters such that objects are grouped in the same cluster when they are “similar†according to similarities, traits and behavior. The effectiveness and efficiency of the existing algorithms is, somewhat limited, since clustering with big data requires clustering high-dimensional feature vectors and since big data often contain large amounts of noise and large datasets. K-Means and DBSCAN are complement to analyze big data on cloud environment. A hybrid approach based on parallel K-Means and parallel DBSCAN is proposed to overcome the drawbacks of both these algorithms. The hybrid approach combines the benefits of both the clustering techniques. The proposed technique is evaluated on the MapReduce framework of Hadoop Platform. The results show that the proposed hybrid approach is an improved version of parallel K-Means clustering algorithm and parallel DBSCAN algorithm. |
Other Details |
|
Paper ID: IJSRDV3I30504 Published in: Volume : 3, Issue : 3 Publication Date: 01/06/2015 Page(s): 3594-3598 |
Article Preview |
|
|
|
|
