Data Extraction and Annotation Methods Using Tag Value Structure
Journal: International Journal of Science and Research (IJSR) (Vol.4, No. 7)Publication Date: 2015-07-05
Authors : Tushar Jadhav; Santosh Chobe;
Page : 1968-1972
Keywords : data extraction; data annotation; data alignment; wrapper generation;
Abstract
The world wide web generates search result pages which is based. On the users input query. It is very crucial for many applications like data integration which requires combining more databases to automatically extract the data from the search results. A unique method for extracting the data and then aligning is implemented which uses Unsupervised duplicate detection algorithm which identifies and segments the result records first and then aligns the segmented results in a table, in which data values of similar attributes are put in same column. The new technique is implemented so as to handle the case when the search results are not adjoining which might happen because of auxiliary data such as advertisements, comments etc. and also to handle nested tag structure which might be present in the search results. The results shows that the implemented algorithm performs well than existing methods.
Other Latest Articles
- Analysis of Radicular Dental Calcifications According to the Shape
- Anthropometric Relation between Height and Arm Length in Adult Male Population of Faridabad, Haryana
- Study The Electronic Transition Behavior of Divalent Transition Metal Cations Co (II), Ni(II) and Cu(II) in Aluminum Nitrate/Urea Room Temperature Ionic Liquid (1:1.2)
- Fuzzy Verdict Reveal with Pruning and Rule Extraction
- Study on Search Behaviours Using Task Trail
Last modified: 2021-06-30 21:50:52