Analysis of Different Similarity Measures for Text Document Comparison |
Author(s): |
| R. Anushya , NGM College, pollachi; Dr. Antony Selvadoss Thanamani, NGM College, Pollachi |
Keywords: |
| Document Similarity, Jaccard Similarity, Cosine Similarity, Euclidean Distance, Pearson Coefficient |
Abstract |
|
Present days people are related with substantial measure of information on standard premise. The sole reason for produced information is to meet the quick needs and no endeavor in sorting out the information for later proficient recovery. Data mining is an idea of separating learning from such a gigantic measure of information. While a few grouping techniques and the related similitude measures have been proposed before, there is no deliberate near investigation of the effect of closeness measures on bunch quality. This might be on the grounds that the famous cost criteria don't promptly decipher crosswise over subjectively extraordinary measures. In this paper we look at four well known likeness measures (Euclidean, cosine, Pearson relationship and broadened Jaccard) related to various kinds of vector space portrayal (Boolean, term recurrence and term recurrence and reverse report recurrence) of archives. Grouping of reports is performed utilizing summed up k-Means; a Partitioned put together bunching strategy with respect to high dimensional inadequate information speaking to content records. Execution is estimated against a human-forced grouping of Topic and Place classifications. We led various analyses and utilized entropy measure to guarantee factual essentialness of results. Cosine, Pearson relationship and broadened Jaccard similitudes rise as the best measures to catch human classification conduct, while Euclidean measures perform poor. |
Other Details |
|
Paper ID: IJSRDV6I100264 Published in: Volume : 6, Issue : 10 Publication Date: 01/01/2019 Page(s): 776-780 |
Article Preview |
|
|
|
|
