Survey on Optimized Feature Subset Selection for the High Dimensional Data |
Author(s): |
| Vidyashree Namdeo Alone , Vidyalankar Institute of Technology, Wadala, Dadar(E),Mumbai.; Prof. Vidya Chitre, Vidyalankar Institute of Technology, Wadala, Dadar(E),Mumbai. |
Keywords: |
| Irrelevant and Redundant Features, Fast Clustering-Based Feature Selection Algorithm, Feature Subset Selection, Data Mining, Filter Method, Featured Clustering, Data Search, Clustering, Rule Mining |
Abstract |
|
Feature selection involves identifying a subset of the most useful features that results having similar weightage as the original entire set of features. A feature selection algorithm may be find from varied points of view but its performance in terms of time complexity has the most eminent as its efficiency concerns the time required to find a subset of features, the efficiency has been related to the quality of the subset of features. Based on this criteria, a clustering-based feature selection algorithm is proposed here. The algorithm contains in two steps. In the first step, features are partitioned into clusters by using cosine analysis over the large multidimensional set of data stored in the system related to various sections of medical data. In the second step, the most representative feature that has strongly related to each cluster (or class) has been selected from each cluster to form a subset of features. Features in distinct clusters have relatively independent; the clustering-based strategy of proposed algorithm has a high probability of producing a subset of useful and independent features. To ensure the efficiency of proposed algorithm, we adopt the efficient Kruskal minimum spanning tree (MST) clustering method. The time complexity of the proposed algorithm is evaluated here. |
Other Details |
|
Paper ID: IJSRDV3I50433 Published in: Volume : 3, Issue : 5 Publication Date: 01/08/2015 Page(s): 436-439 |
Article Preview |
|
|
|
|
