Web Data Extraction and Alignment
International Journal of Science and Research (IJSR)

International Journal of Science and Research (IJSR)
Call for Papers | Fully Refereed | Open Access | Double Blind Peer Reviewed

ISSN: 2319-7064


Downloads: 121 | Views: 385

Research Paper | Computer Science & Engineering | India | Volume 2 Issue 3, March 2013 | Popularity: 6.9 / 10


     

Web Data Extraction and Alignment

M. Jude Victor, D. John Aravindhar, V. Dheepa


Abstract: Web databases generate query result pages based on a user’s query. Automatically extracting the data from these query result pages is very important for many applications, such as data integration, which need to cooperate with multiple web databases. We present a novel data extraction and alignment method called CTVS that combines both tag and value similarity. CTVS automatically extracts data from query result pages by first identifying and segmenting the query result records (QRRs) in the query result pages and then aligning the segmented QRRs into a table, in which the data values from the same attribute are put into the same column. We also design a new record alignment algorithm that aligns the attributes in a record, first pair wise and then holistically, by combining the tag and data value similarity information. Experimental results show that CTVS achieves high precision and outperforms existing state-of-the-art data extraction methods.


Keywords: Data Extraction, Automatic Wrapper Generation, Data Record Alignment, Information Integration


Edition: Volume 2 Issue 3, March 2013


Pages: 129 - 132



Please Disable the Pop-Up Blocker of Web Browser

Verification Code will appear in 2 Seconds ... Wait



Text copied to Clipboard!
M. Jude Victor, D. John Aravindhar, V. Dheepa, "Web Data Extraction and Alignment", International Journal of Science and Research (IJSR), Volume 2 Issue 3, March 2013, pp. 129-132, https://www.ijsr.net/getabstract.php?paperid=IJSROFF2013098, DOI: https://www.doi.org/10.21275/IJSROFF2013098

Top