Document Clustering using Enhanced Similarity Measurement for Text Processing |
Author(s): |
| M.Krishnamoorthy , K.S.R college of engineering; V.Sharmila, K.S.R college of engineering; P.Balamurugan, K.S.R college of engineering |
Keywords: |
| Document classification, document clustering, entropy, accuracy, classifiers, clustering algorithms |
Abstract |
|
Clustering is one of the most important techniques in machine learning and data mining tasks. Similar data grouping is performed using clustering techniques. In document vector each component indicates the value of the corresponding feature in the document. The feature value can be term frequency, relative term frequency. Similarity Measurement for Text Process (SMTP) is used to compute the similarity between two documents with respect to a feature. Presents and options of the features in both documents are used to estimate the similarity values. The SMTP is extended to estimate similarity between two set of documents. The SMTP scheme is used with text clustering and classification task.The system is designed to perform document clustering using Similarity Measurement for Text Process (SMTP). Spherical K means algorithm is used for the clustering process. Concept relationships are identified with the support of ontology. Dimensionality reduction functions are applied to minimize the features. |
Other Details |
|
Paper ID: IJSRDV2I12131 Published in: Volume : 2, Issue : 12 Publication Date: 01/03/2015 Page(s): 280-285 |
Article Preview |
|
|
|
|
