IJIRST (International Journal for Innovative Research in Science & Technology)ISSN (online) : 2349-6010

 International Journal for Innovative Research in Science & Technology

A Hybrid Unsupervised Web Data Extraction using Trinity and NLP


Print Email Cite
International Journal for Innovative Research in Science & Technology
Volume 2 Issue - 2
Year of Publication : 2015
Authors : Anju R ; Mary Mareena P V

BibTeX:

@article{IJIRSTV2I2073,
     title={A Hybrid Unsupervised Web Data Extraction using Trinity and NLP},
     author={Anju R and Mary Mareena P V },
     journal={International Journal for Innovative Research in Science & Technology},
     volume={2},
     number={2},
     pages={229--233},
     year={},
     url={http://www.ijirst.org/articles/IJIRSTV2I2073.pdf},
     publisher={IJIRST (International Journal for Innovative Research in Science & Technology)},
}



Abstract:

Web is a huge repository of data. Inorder to automatically extract relevant data from web documents, web data extractors are used. The proposed technique works on two web documents that are generated by the same server-side template and learns a regular expression which represents the template of the web document. The regular expression generated can be later used to extract data from other similar documents. The proposed technique builds on the hypothesis that template introduces some shared pattern that do not provide any relevant data. In the regular expression the capturing groups represent the data. The semantic label for each capturing group is provided based on the POS (Part Of Speech) tagging. This technique could even work with the malformed documents. Hence the input errors do not have negative impact on the effectiveness.


Keywords:

Web data extraction, Wrapper induction, Unsupervised Techniques, NLP, POS tagging


Download Article