High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Big Data and Hadoop: Improving Mapreduce Performance using Clustering Algorithm

Author(s):

Krunal Dave , L.j Institute of engineering and technology; Mr.Jignesh Vania, L.j Institute of engineering and technology

Keywords:

Big data, Clustering, Hadoop, k-means, MapReduce, HDFS

Abstract

“Big Data” is a popular term used to describe the exponential growth and availability of structured, unstructured data and semi-structured data that has potential to be mined for information. Data mining involves knowledge discovery from these large data sets. Hadoop is the core platform for storing the large volume of data into Hadoop Distributed File System (HDFS) and that data get processed by MapReduce model in parallel. Hadoop is designed to scale up from a single server to thousands of machines and with a very high degree of fault tolerance. This paper presents we have implemented our k-means algorithm in single and multi-node Hadoop cluster. We calculate the performance based on the time to execute the MapReduce job.

Other Details

Paper ID: IJSRDV3I30951
Published in: Volume : 3, Issue : 3
Publication Date: 01/06/2015
Page(s): 3571-3574

Article Preview

Download Article