An Efficient Pre-processing Mechanism for Web Usage Mining |
Author(s): |
| Dipika Sahu , shri shankaracharya group of institutions ; Yamini Chouhan, shri shankaracharya group of institutions |
Keywords: |
| Web Mining, Data Cleaning, Session Identification, Data Formatting and Log Querying |
Abstract |
|
Web mining is the application of data mining techniques to extract knowledge from Web data, including Web document, hyperlink between document, usages logs of web sites, etc. Data preprocessing has a fundamental role in Web Usage Mining applications. There are two stages [1] Data clean consist of removing all the data records in web logs that are useless for mining purpose e.g.: request for graphical page content (e.g., jpg and gif images); requests for any other file which may be include into a web page; or even navigation sessions performed by robots and web spiders. [2] Session Identification and Reconstruction consists set of pages visited by the same user within the duration of one particular visit to a web site. Data Formatting is the final step of preprocessing. Once the previous phases have been complete, data are properly formatted before applying mining techniques, stores data extracted from web logs into a relational database using a click fact scheme, so as to provide superior support to log querying finalized to frequent pattern mining. |
Other Details |
|
Paper ID: IJSRDV4I80343 Published in: Volume : 4, Issue : 8 Publication Date: 01/11/2016 Page(s): 763-767 |
Article Preview |
|
|
|
|
