Improving Data Locality In Hadoop Framework For Big data applications
This Project has made an attempt to show background of massive information management and referenced numerous issues and difficulties faced to process large amount of information. Many execution based limitations are observed while Hadoop MapReduce is utilized to execute jobs which require prom
2025-06-28 16:33:05 - Adil Khan
Improving Data Locality In Hadoop Framework For Big data applications
Project Area of Specialization Wearables and ImplantableProject SummaryThis Project has made an attempt to show background of massive information management and referenced numerous issues and difficulties faced to process large amount of information. Many execution based limitations are observed while Hadoop MapReduce is utilized to execute jobs which require prompt reactions; also, potential solutions are examined and presented. As to increase the execution of map and reduce jobs, enhancement methodologies are needed to increase the execution of MapReduce performance to enhance data locality. A better scheduling algorithm has been considered that is iShuffle which can result in great execution and outputs. Then we examined some limitations from ishuffle algorithm and purposed a new solution Eishuffle based on job size, partition size and enhance real time processing. EIShuffle gives better results and increase the CPU utilization and decrease the responce time.
Project ObjectivesThe following are the goals of research project:
- In data locality, the program is moved to the node where data resides instead of moving large data to the node, this makes Hadoop faster. Because the size of the program is always lesser than the size of data, so moving data is a bottleneck of network transfer.
- Try to minimize network congestion and increase the overall throughput of the system.
- To improve reliability and reduce cost of data accessibility.
- To help the end users to access data faster and reduce the execution time in Hadoop.
- Try to review all previous research on data locality in Hadoop and improve it.
A new solution is proposed in the project and is implemented in 4-node cloudera Hadoop on google.
The Proposd solution is better than the existing Techniques.
Benefits of the ProjectThe main purpose of this project is basically to improve data locaity In hadoop so that the management of Big Data would be Easy. In oyr work we deal with the responce time and real time processing og the Data Jobs.
Technical Details of Final DeliverableA Research paper named " A critical analysis:Improving Data locality in Hadoop Framwork for Big Data Applications " is written and Published by INCCST Conference.
Another Research paper is in progress in which the testing and results of the proposed solution are discussed.
Final Deliverable of the Project Software SystemType of Industry IT Technologies Big DataSustainable Development Goals Quality EducationRequired Resources| Item Name | Type | No. of Units | Per Unit Cost (in Rs) | Total (in Rs) |
|---|---|---|---|---|
| Total in (Rs) | 70000 | |||
| Publication Fee | Miscellaneous | 1 | 5000 | 5000 |
| Domain and Hosting | Equipment | 3 | 20000 | 60000 |
| Survey Travel Cost | Miscellaneous | 1 | 5000 | 5000 |