High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Big Data Processing: Improve Scheduling Environment in Hadoop

Author(s):

Joshi Bhavik Bipinkumar , LDRP-ITR

Keywords:

Big Data, Hadoop, Map Reduce, Scheduling algorithms, SARS algorithm

Abstract

Now a days, dealing with datasets in the order of terabytes or even petabyte is a reality. Therefore, processing such big datasets in an efficient way is a clear need for many users. In this context, Hadoop Map Reduce is a big data processing framework that has rapidly become the de facto standard in both industry and academia. The main reasons of such popularity are the ease-of-use, scalability, and failover properties of Hadoop Map Reduce. However, these features come at a price: the performance of Hadoop Map Reduce is usually far from the performance of a well-tuned parallel database. Therefore, many research works (from industry and academia) have focused on improving the performance of Hadoop Map Reduce jobs in many aspects. For example, researchers have proposed different data layouts; join algorithms, high-level query languages, failover algorithms, query optimization technique, and indexing techniques. We will point out the similarities and differences between the techniques used in Hadoop with those used in parallel databases. And try to improving scheduling environment in Hadoop.

Other Details

Paper ID: IJSRDV4I60236
Published in: Volume : 4, Issue : 6
Publication Date: 01/09/2016
Page(s): 519-524

Article Preview

Download Article