High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Text Categorization on Multiple Languages based on Classification Technique

Author(s):

Bhavna Rani , department of computer science and applications kurukshetra university kurukshetra

Keywords:

Vector Space Model, ImproveKNN(K-NearestNeighbour), precision(p),Recall(r), F-measure, tokens, Stemming, Stopwords

Abstract

In the Constitution of India, a provision is made for each of the Indian states to choose their own official language for communicating at the state level for official purpose. The availability of constantly increasing amount of textual data of various Indian regional languages in electronic form has accelerated. So the Classification of text documents based on languages is essential. The objective of the work is the representation and categorization of Indian language text documents using text mining techniques. Several text mining techniques such as Vector Space Model,ImproveKNN(KNearestNeighbour),precision(p),Recall(r),Fmeasure,tokens,Stemming,Stopwords for text categorization have been used.

Other Details

Paper ID: IJSRDV3I31564
Published in: Volume : 3, Issue : 3
Publication Date: 01/06/2015
Page(s): 3522-3524

Article Preview

Download Article