High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Extraction of Query Results from Deep Web Interfaces Automatically

Author(s):

Abhishek Popat Bangar , SPCOE,Dumbarwadi; Prof. Kore Kunal S., SPCOE,Dumbarwadi

Keywords:

Extraction, Deep Web, Parallel Crawler, Focused Web Crawler, Web Crawler

Abstract

The content unseen behind HTML forms, has shorty been documented as a significant gap in search engine coverage. It represents vital contents of the data on the Web; retrieving Deep-Web content is not an easy task for the database community. Indexing of the searched data is major problem tackled by web crawlers that has deeply effect on search engine efficiency. Latest study about searching contents on the web illustrate that nearly 96% of data over internet is encapsulated as well as hidden i.e. hidden from search engines. The main task faced by the search engines is to retrieve and access hidden web data (web interfaces) or contents at low cost. The proposed system uses a machine learning approach that is very scalable, totally automatic, and identically efficient to use, that helps to expand data retrieval functionality at lower cost. The proposed system uses focused crawling strategy for accessing perfect searched results related to query and pick out only relevant information or data according to their similarity with respect to query. The proposed algorithm can selects only possible candidates rather than searching whole document for addition in to your web search index. The automatic attribute building is used for form classification that helps to reduce manual training time and data set building.

Other Details

Paper ID: IJSRDV6I50259
Published in: Volume : 6, Issue : 5
Publication Date: 01/08/2018
Page(s): 553-555

Article Preview

Download Article