A Two-Stage Smart Crawler for Deep Web-Page Search |
Author(s): |
| Zaid Pathan , Dr D Y Patil Institute of Technology, Pimpri; Shreyas Reddy, Dr D Y Patil Institute of Technology, Pimpri; Taha Rizvi, Dr D Y Patil Institute of Technology, Pimpri; Rahul Jain, Dr D Y Patil Institute of Technology, Pimpri |
Keywords: |
| Deep Web Harvesting, Focused Crawler, Ranking URLs, Data Sharing, Search Interface Discovery, Green Computing |
Abstract |
|
The Internet is a global, common and self-sufficient structure available to billions of people around the world. Huge amount of data is stored in a structured format. When the user looks for data in the search engine, it returns huge volume of relevant and irrelevant data. The returned queries are arranged with the help of certain parameters, wherein number of visits being one of them. Due to this, the new sites having higher relevancy, but less number of visits suffer and hence are ranked low. This is referred as cold start problem. To overcome this problem, we propose a two-stage Smart Crawler, to extract the relevant data from the Invisible Web. This module returns seed sites from the site database. In the first stage, this framework performs a 'reverse lookup' that matches the users’ query with URL. In the second stage, the 'prioritization of incremental sites' is performed as per the content of the query. The returned links are classified as relevant and irrelevant links, and high ranked links are shown on the results page. Our proposed model efficiently returns deep web interfaces from a pool of sites and achieves better results than other crawlers, as well as eliminates the cold start problem by returning only relevant data. To achieve performance, we design 'custom search' where the user gets data based on their profession. |
Other Details |
|
Paper ID: IJSRDV5I110220 Published in: Volume : 5, Issue : 11 Publication Date: 01/02/2018 Page(s): 446-449 |
Article Preview |
|
|
|
|
