Large Scales Spam Emails Filtering with Mapreduce based SVM |
Author(s): |
| Dipika Somvanshi , Bharati Vidyapeeth's College Of Engineering; Prof. Kanchan Doke, Bharati Vidyapeeth's College Of Engineering |
Keywords: |
| Spam Filtering, Machine Learning Techniques, Naïve Bays, KNN, Decision Trees, SVM, MapReduce |
Abstract |
|
Spam is any unwanted and harmful mail send to massive recipients in bulk quantity. Spam can be harmful as it may contain malware and links to phishing websites or harmless as advertisement promotion content. The volume of spam has been increasing significantly over last few decades and therefore there is an urgent need to address the Email spam problem. Machine learning techniques are most popular because of high accuracy and mathematical support. SVM is the popular machine learning techniques in spam filtering because it can handle data with large number of attributes. MapReduce has become increasingly popular as an Internet scale, data intensive processing platform. In the context of machine learning based spam filter training, support vector machine (SVM) based techniques have been proven effective. SVM training is however a computationally intensive process. SVM can’t handle large dataset as input. These drawbacks of SVM are overcome by MapReduce. In this dissertation, a MapReduce based distributed SVM algorithm for large dataset of E-mails spam filter training, is proposed. It gives large scalability and speedup the performance of spam filter efficiently. |
Other Details |
|
Paper ID: IJSRDV5I50884 Published in: Volume : 5, Issue : 5 Publication Date: 01/08/2017 Page(s): 1160-1163 |
Article Preview |
|
|
|
|
