Evaluating Classification Techniques on Different Categories of Datasets |
Author(s): |
| Sachin Singh Thakur , IIIT JABALPUR; Shubham Prakash Srivastava, UNIVERSITY OF PETROLEUM AND ENERGY STUDIES |
Keywords: |
| Statistical Technique, Machine Learning Technique, Area under Curve, Imbalanced Datasets, Missing Value Datasets, Noisy Datasets, Size Datasets |
Abstract |
|
This paper aims to address most popular and interesting problem from the field of classification. Choosing a suitable classifier for a given problem which itself has been a challenging problem that typically requires experimentation with different classification techniques. In our work, we try to address this question by investigating any relation that may exist between the dataset characteristics and classifier bias. Specifically, in our study, we categorize datasets based on their inherent Characteristics like imbalanced, missing value, noise, and different sizes of the datasets and investigate the performance of two classes of classifiers - five Machine Learning and five Statistical Learning classifiers. Machine Learning techniques scored over Statistical techniques in five categories out of seven categories of the datasets. Multilayer Perceptron was found to be the most robust classifiers in the class of machine learning classifiers and Bayes Net was found to be the most robust classifiers in the class of Statistical Learning classifiers. |
Other Details |
|
Paper ID: IJSRDV6I20166 Published in: Volume : 6, Issue : 2 Publication Date: 01/05/2018 Page(s): 3821-3831 |
Article Preview |
|
|
|
|
