Text Categorization on Multiple Languages based on Classification Technique |
Author(s): |
| Bhavna Rani , department of computer science and applications kurukshetra university kurukshetra |
Keywords: |
| Vector Space Model, ImproveKNN(K-NearestNeighbour), precision(p),Recall(r), F-measure, tokens, Stemming, Stopwords |
Abstract |
|
In the Constitution of India, a provision is made for each of the Indian states to choose their own official language for communicating at the state level for official purpose. The availability of constantly increasing amount of textual data of various Indian regional languages in electronic form has accelerated. So the Classification of text documents based on languages is essential. The objective of the work is the representation and categorization of Indian language text documents using text mining techniques. Several text mining techniques such as Vector Space Model,ImproveKNN(KNearestNeighbour),precision(p),Recall(r),Fmeasure,tokens,Stemming,Stopwords for text categorization have been used. |
Other Details |
|
Paper ID: IJSRDV3I31564 Published in: Volume : 3, Issue : 3 Publication Date: 01/06/2015 Page(s): 3522-3524 |
Article Preview |
|
|
|
|
