• [email protected]
  • +971 507 888 742
Submit Manuscript
SciAlert
  • Home
  • Journals
  • Information
    • For Authors
    • For Referees
    • For Librarian
    • For Societies
  • Contact
  1. Information Technology Journal
  2. Vol 11 (4), 2012
  3. 408-413
  • Issues
    Online First Current Issue All Issues
  • Information About
    Aims and Scope Editorial Board Guide to Authors Article Processing Charges
    Submit a Manuscript

Information Technology Journal

Year: 2012 | Volume: 11 | Issue: 4 | Page No.: 408-413
DOI: 10.3923/itj.2012.408.413

Facebook Twitter Reddit Linkedin E-mail
Google Scholar ASCI
Research Article

Web Information Extraction Based on Visual Characteristics

Ruyue Tan
Room243 Apartment7, Harbin Institute of Technology, Wenhuaxilu 2, Huancuiqu, Weihai, Shandong, People`s Republic of China

Due to the explosive development of Internet technology, the web is becoming the world's largest database of information, effective management and utilization of web information is currently a hot issue. This study mainly discusses the extraction of web information. Traditional web information extraction is mainly based on DOM tree and HTML tag analysis. Based on VIPS, the study proposes visual block positioning algorithm for webpage information extraction through induction webpage visual features and visual block feature information. It inputs the theme-based webpages and BBS webpages for VIPS, analyzes the output of VIPS and the VBT and then defines visual characteristics such as text density and link text density. The study puts forward a visual block positioning algorithm VBPA. It will position the theme information block to one VBT node, then extract the theme information. Experimental results show that the visual block positioning algorithm based on visual features is superior to the traditional web information extraction algorithm and has a higher quality of information extraction.
PDF Fulltext XML References Citation

How to cite this article

Ruyue Tan, 2012. Web Information Extraction Based on Visual Characteristics. Information Technology Journal, 11: 408-413.

DOI: 10.3923/itj.2012.408.413

URL: https://scialert.net/abstract/?doi=itj.2012.408.413

Leave a Comment


Your email address will not be published. Required fields are marked *

Article Trend



Total views 1842

References


  1. Cai, D., S. Yu, J.R. Wen and W.Y. Ma, 2003. Extracting content structure for web pages based on visual representation. Proceedings of the 5th Asia Pacific Web Conference, April 23-25, 2003, Xi'an China, pp: 406-417.
    CrossRef

  2. Cai, D., S. Yu, J.R. Wen and W.Y. Ma, 2003. VIPS: A vision-based page segmentation algorithm. Microsoft Technical Report, MSR-TR-203-79. http://research.microsoft.com/apps/pubs/default.aspx?id=70027.

  3. Liu, W. and X. Meng, 2006. Vision-based web and data records extraetion. Proceedings of the 9th SIGMOD International Workshop on Web and Databases, June 30, 2006, Chicago.

  4. Chang, C.H., M. Kayed, R. Girgis and K.F. Shaalan, 2006. A survey of web information extraction system. Inst. Electr. Electron. Eng. Trans. Knowledge Data Eng., 18: 1411-1428.
    CrossRef

  5. Sarawagi, S., 2002. Automation in information extraction and integration. Proceedings of the 28th International Conference on Very Large Data Bases, August 20-23, 2002, Hong Kong, China.

  6. Laender, A.H.F., B.A. Ribeiro-Neto, A.S. da Silva and J.S. Teixeira, 2002. A brief survey of web data extraction tools. ACM SIGMOD Record, 31: 84-93.
    CrossRefDirect Link

  7. Crescenzi, V. and G. Mecca, 1998. Grammars have exceptions. Inform. Syst., 23: 539-565.
    CrossRef

  8. Saiiuguet, A. and F. Azavant, 2001. Building intelligent web applications using lightweight wrappers. Data Knowledge Eng., 36: 283-316.
    Direct Link

Keywords


  • Subject extraction
  • BBS information extraction
  • VIPS
  • visual pieces positioning; VBPA

Useful Links

  • Journals
  • For Authors
  • For Referees
  • For Librarian
  • For Socities

Contact Us

Office Number 1128,
Tamani Arts Building,
Business Bay,
Deira, Dubai, UAE

Phone: +971 507 888 742
Email: [email protected]

About Science Alert

Science Alert is a technology platform and service provider for scholarly publishers, helping them to publish and distribute their content online. We provide a range of services, including hosting, design, and digital marketing, as well as analytics and other tools to help publishers understand their audience and optimize their content. Science Alert works with a wide variety of publishers, including academic societies, universities, and commercial publishers.

Follow Us
© Copyright Science Alert. All Rights Reserved