<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "http://dtd.nlm.nih.gov/publishing/2.0/journalpublishing.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" article-type="research-article" dtd-version="2.0" xml:lang="EN">

  <front>

    <journal-meta>

      <journal-title>Information Technology Journal</journal-title>

      <issn pub-type="ppub">1812-5638</issn>

      <issn pub-type="epub">1812-5646</issn>

      <publisher>

        <publisher-name>Asian Network for Scientific Information</publisher-name>

      </publisher>

    </journal-meta>


    <article-meta>

      <article-id pub-id-type="doi">10.3923/itj.2012.408.413</article-id>


      <title-group>

        <article-title><![CDATA[Web Information Extraction Based on Visual Characteristics]]></article-title>

      </title-group>


      <contrib-group>

        <contrib contrib-type="author" xlink:type="simple">


          <name name-style="western">

            <surname>Tan</surname>

            <given-names>Ruyue</given-names>

          </name>


        </contrib>

      </contrib-group>


      <pub-date pub-type="collection">


        <month>4</month>




        <year>2012</year>

      </pub-date>


      <volume>11</volume>

      <issue>4</issue>


      <abstract><![CDATA[<p>Due to the explosive development of Internet technology, the web is becoming the world's largest database of information, effective management and utilization of web information is currently a hot issue. This study mainly discusses the extraction of web information. Traditional web information extraction is mainly based on DOM tree and HTML tag analysis. Based on VIPS, the study proposes visual block positioning algorithm for webpage information extraction through induction webpage visual features and visual block feature information. It inputs the theme-based webpages and BBS webpages for VIPS, analyzes the output of VIPS and the VBT and then defines visual characteristics such as text density and link text density. The study puts forward a visual block positioning algorithm VBPA. It will position the theme information block to one VBT node, then extract the theme information. Experimental results show that the visual block positioning algorithm based on visual features is superior to the traditional web information extraction algorithm and has a higher quality of information extraction.</p>]]></abstract>


    </article-meta>

  </front>


  <ref-list>










      <ref id="4304">

        <label>1</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Cai, D., S. Yu, J.R. Wen and W.Y. Ma,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2003</year>

          <article-title><![CDATA[Extracting content structure for web pages based on visual representation.]]></article-title>

          <source>Proceedings of the 5th Asia Pacific Web Conference,</source>

          <volume>2003</volume>

          <fpage>pp: 406</fpage>

          <lpage>417</lpage>

        </citation>

      </ref>




















      <ref id="45412">

        <label>2</label>

        <citation citation-type="anonymous" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Cai, D., S. Yu, J.R. Wen and W.Y. Ma,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2003</year>

          <article-title><![CDATA[VIPS: A vision-based page segmentation algorithm.]]></article-title>

          <source>VIPS: A vision-based page segmentation algorithm.</source>

          <volume>2003</volume>

        </citation>

      </ref>
















      <ref id="34495">

        <label>3</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Liu, W. and X. Meng,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2006</year>

          <article-title><![CDATA[Vision-based web and data records extraetion.]]></article-title>

          <source>Proceedings of the 9th SIGMOD International Workshop on Web and Databases,</source>

          <volume>2006</volume>

          <fpage></fpage>

          <lpage></lpage>

        </citation>

      </ref>












      <ref id="854958">

        <label>4</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Chang, C.H., M. Kayed, R. Girgis and K.F. Shaalan,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2006</year>

          <article-title><![CDATA[A survey of web information extraction system.]]></article-title>

          <source>Inst. Electr. Electron. Eng. Trans. Knowledge Data Eng.,</source>

          <volume>18</volume>

          <fpage>1411</fpage>

          <lpage>1428</lpage>

        </citation>

      </ref>
























      <ref id="34500">

        <label>5</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Sarawagi, S.,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2002</year>

          <article-title><![CDATA[Automation in information extraction and integration.]]></article-title>

          <source>Proceedings of the 28th International Conference on Very Large Data Bases,</source>

          <volume>2002</volume>

          <fpage></fpage>

          <lpage></lpage>

        </citation>

      </ref>












      <ref id="34628">

        <label>6</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Laender, A.H.F., B.A. Ribeiro-Neto, A.S. da Silva and J.S. Teixeira,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2002</year>

          <article-title><![CDATA[A brief survey of web data extraction tools.]]></article-title>

          <source>ACM SIGMOD Record,</source>

          <volume>31</volume>

          <fpage>84</fpage>

          <lpage>93</lpage>

        </citation>

      </ref>


















      <ref id="854976">

        <label>7</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Crescenzi, V. and G. Mecca,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1998</year>

          <article-title><![CDATA[Grammars have exceptions.]]></article-title>

          <source>Inform. Syst.,</source>

          <volume>23</volume>

          <fpage>539</fpage>

          <lpage>565</lpage>

        </citation>

      </ref>


















      <ref id="854982">

        <label>8</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Saiiuguet, A. and F. Azavant,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2001</year>

          <article-title><![CDATA[Building intelligent web applications using lightweight wrappers.]]></article-title>

          <source>Data Knowledge Eng.,</source>

          <volume>36</volume>

          <fpage>283</fpage>

          <lpage>316</lpage>

        </citation>

      </ref>
















  </ref-list>

</article>

