Big Data Processing: Improve Scheduling Environment in Hadoop |
Author(s): |
| Joshi Bhavik Bipinkumar , LDRP-ITR |
Keywords: |
| Big Data, Hadoop, Map Reduce, Scheduling algorithms, SARS algorithm |
Abstract |
|
Now a days, dealing with datasets in the order of terabytes or even petabyte is a reality. Therefore, processing such big datasets in an efficient way is a clear need for many users. In this context, Hadoop Map Reduce is a big data processing framework that has rapidly become the de facto standard in both industry and academia. The main reasons of such popularity are the ease-of-use, scalability, and failover properties of Hadoop Map Reduce. However, these features come at a price: the performance of Hadoop Map Reduce is usually far from the performance of a well-tuned parallel database. Therefore, many research works (from industry and academia) have focused on improving the performance of Hadoop Map Reduce jobs in many aspects. For example, researchers have proposed different data layouts; join algorithms, high-level query languages, failover algorithms, query optimization technique, and indexing techniques. We will point out the similarities and differences between the techniques used in Hadoop with those used in parallel databases. And try to improving scheduling environment in Hadoop. |
Other Details |
|
Paper ID: IJSRDV4I60236 Published in: Volume : 4, Issue : 6 Publication Date: 01/09/2016 Page(s): 519-524 |
Article Preview |
|
|
|
|
