High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

An Efficient Pre-processing Mechanism for Web Usage Mining

Author(s):

Dipika Sahu , shri shankaracharya group of institutions ; Yamini Chouhan, shri shankaracharya group of institutions

Keywords:

Web Mining, Data Cleaning, Session Identification, Data Formatting and Log Querying

Abstract

Web mining is the application of data mining techniques to extract knowledge from Web data, including Web document, hyperlink between document, usages logs of web sites, etc. Data preprocessing has a fundamental role in Web Usage Mining applications. There are two stages [1] Data clean consist of removing all the data records in web logs that are useless for mining purpose e.g.: request for graphical page content (e.g., jpg and gif images); requests for any other file which may be include into a web page; or even navigation sessions performed by robots and web spiders. [2] Session Identification and Reconstruction consists set of pages visited by the same user within the duration of one particular visit to a web site. Data Formatting is the final step of preprocessing. Once the previous phases have been complete, data are properly formatted before applying mining techniques, stores data extracted from web logs into a relational database using a click fact scheme, so as to provide superior support to log querying finalized to frequent pattern mining.

Other Details

Paper ID: IJSRDV4I80343
Published in: Volume : 4, Issue : 8
Publication Date: 01/11/2016
Page(s): 763-767

Article Preview

Download Article