High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Clustering Based Feature Subset Selection for High Dimensional data

Author(s):

Lalith Prasad K A , RNSIT; Prakasha S, RNSIT

Keywords:

Clustering , Subset Selection , High Dimensional data, Soft Computing.

Abstract

Feature selection is a term commonly used in data mining to describe the tools and techniques available for reducing inputs to a manageable size for processing and analysis. Feature selection implies not only cardinality reduction, which means imposing an arbitrary or predefined cutoff on the number of attributes that can be considered when building a model, but also the choice of attributes, meaning that either the analyst or the modeling tool actively selects or discards attributes based on their usefulness for analysis. Feature selection involves identifying a subset of the most useful features that produces compatible results as the original entire set of features. A feature selection algorithm may be evaluated from both the efficiency and effectiveness points of view. While the efficiency concerns the time required to find a subset of features, the effectiveness is related to the quality of the subset of features. The proposed algorithm works in two steps. In the first step, features are divided into clusters by using graph-theoretic clustering methods. In the second step, the most representative feature that is strongly related to target classes is selected from each cluster to form a subset of features. Features in different clusters are relatively independent; the clustering-based strategy of this algorithm has a high probability of producing a subset of useful and independent features. To ensure the efficiency of this algorithm, we adopt the efficient minimum-spanning tree clustering method. Features in different clusters are relatively independent; the clustering-based strategy of this algorithm has a high probability of producing a subset of useful and independent features. The efficiency and the effectiveness of the algorithm can be evaluated by the runtime efficiency, classification accuracy and the subset of features selected. Extensive experiments are performed on this algorithm by comparing it with known classifiers like tree-based C4.5. The results on the high dimensional image and microarray data, demonstrate that the algorithm not only produces smaller subsets of features but also improves the performance of the C4.5 classifier.

Other Details

Paper ID: IJSRDV2I4194
Published in: Volume : 2, Issue : 4
Publication Date: 01/07/2014
Page(s): 316-319

Article Preview

Download Article