High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Analysis of Different Similarity Measures for Text Document Comparison

Author(s):

R. Anushya , NGM College, pollachi; Dr. Antony Selvadoss Thanamani, NGM College, Pollachi

Keywords:

Document Similarity, Jaccard Similarity, Cosine Similarity, Euclidean Distance, Pearson Coefficient

Abstract

Present days people are related with substantial measure of information on standard premise. The sole reason for produced information is to meet the quick needs and no endeavor in sorting out the information for later proficient recovery. Data mining is an idea of separating learning from such a gigantic measure of information. While a few grouping techniques and the related similitude measures have been proposed before, there is no deliberate near investigation of the effect of closeness measures on bunch quality. This might be on the grounds that the famous cost criteria don't promptly decipher crosswise over subjectively extraordinary measures. In this paper we look at four well known likeness measures (Euclidean, cosine, Pearson relationship and broadened Jaccard) related to various kinds of vector space portrayal (Boolean, term recurrence and term recurrence and reverse report recurrence) of archives. Grouping of reports is performed utilizing summed up k-Means; a Partitioned put together bunching strategy with respect to high dimensional inadequate information speaking to content records. Execution is estimated against a human-forced grouping of Topic and Place classifications. We led various analyses and utilized entropy measure to guarantee factual essentialness of results. Cosine, Pearson relationship and broadened Jaccard similitudes rise as the best measures to catch human classification conduct, while Euclidean measures perform poor.

Other Details

Paper ID: IJSRDV6I100264
Published in: Volume : 6, Issue : 10
Publication Date: 01/01/2019
Page(s): 776-780

Article Preview

Download Article