<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "http://dtd.nlm.nih.gov/publishing/2.0/journalpublishing.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" article-type="research-article" dtd-version="2.0" xml:lang="EN">

  <front>

    <journal-meta>

      <journal-title>Information Technology Journal</journal-title>

      <issn pub-type="ppub">1812-5638</issn>

      <issn pub-type="epub">1812-5646</issn>

      <publisher>

        <publisher-name>Asian Network for Scientific Information</publisher-name>

      </publisher>

    </journal-meta>


    <article-meta>

      <article-id pub-id-type="doi">10.3923/itj.2013.2447.2453</article-id>


      <title-group>

        <article-title><![CDATA[A New Hybrid Schemes Combining Ontology and Clustering for Text Documents]]></article-title>

      </title-group>


      <contrib-group>

        <contrib contrib-type="author" xlink:type="simple">


          <name name-style="western">

            <surname>Punitha</surname>

            <given-names>S.C.</given-names>

          </name>


          <name name-style="western">

            <surname>Thavavel</surname>

            <given-names>V.</given-names>

          </name>


          <name name-style="western">

            <surname>Punithavalli</surname>

            <given-names>M.</given-names>

          </name>


        </contrib>

      </contrib-group>


      <pub-date pub-type="collection">




        <month>12</month>


        <year>2013</year>

      </pub-date>


      <volume>12</volume>

      <issue>12</issue>


      <abstract><![CDATA[<p>Data mining is a process of analyzing data from different perspectives and summarizing it into valuable information. It consist of two activities such as clustering and classification. It mainly works with numeric data, text data and the web data. Text-based algorithms have problems when dealing with different languages (synonyms, homonyms). Also, web pages contain other forms of information except text, such as images or multimedia. As a consequence, hybrid document clustering approaches have been proposed in order to combine the advantages and limit the disadvantages of the existing approaches. The main motivation behind ontology is that different people have different needs with regard to the clustering of texts. The hybrid schemes are developed using ontology and the frequent item clustering of various algorithms Ontology Based Apriori Based Clustering, Ontology based FP-Growth Based Clustering, Ontology based FP-Bonsai Clustering Algorithm have been proposed to resolve the disadvantages of existing approaches. The performance of this enhanced document clustering algorithm was tested vigorously using different datasets with performance measures to show the efficiency in clustering. Hence Ontology based FP-Bonsai Clustering Algorithm (OFPBC) shows significant improvement in terms of purity of clustering. The result shows that the datasets namely Reuters 21578,20 new Group and TDT2 which results the accuracy 0.840, 0.817 and 0.847 in OFPBC, respectively.</p>]]></abstract>


    </article-meta>

  </front>


  <ref-list>










      <ref id="50193">

        <label>1</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Buckley, C. and A.F. Lewit,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1985</year>

          <article-title><![CDATA[Optimization of inverted vector searches.]]></article-title>

          <source>Proceedings of the 8th Annual International SIGIR Conference on Research and Development in Information Retrieval,</source>

          <volume>1985</volume>

          <fpage>pp: 97</fpage>

          <lpage>110</lpage>

        </citation>

      </ref>












      <ref id="920367">

        <label>2</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Karypis, G., E.H. Han and V.K.P. Kumar, </surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1999</year>

          <article-title><![CDATA[Chameleon: Hierarchical clustering using dynamic modeling.]]></article-title>

          <source>IEEE Comput.,</source>

          <volume>32</volume>

          <fpage>68</fpage>

          <lpage>75</lpage>

        </citation>

      </ref>




















      <ref id="102618">

        <label>3</label>

        <citation citation-type="book" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Euzenat, J. and P. Shvaiko,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2007</year>

          <article-title><![CDATA[Ontology Matching.]]></article-title>

          <source>Ontology Matching.</source>

          <volume>1st Edn.,</volume>

          <fpage>Pages: 333</fpage>

          <lpage>Pages: 333</lpage>

        </citation>

      </ref>


















      <ref id="102619">

        <label>4</label>

        <citation citation-type="book" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Kowalski, G.,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1997</year>

          <article-title><![CDATA[Information Retrieval Systems: Theory and Implementation.]]></article-title>

          <source>Information Retrieval Systems: Theory and Implementation.</source>

          <volume>1st Edn.,</volume>

          <fpage>Pages: 282</fpage>

          <lpage>Pages: 282</lpage>

        </citation>

      </ref>


















      <ref id="11992">

        <label>5</label>

        <citation citation-type="book" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Manning, C. and H. Schutze,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1999</year>

          <article-title><![CDATA[Foundations of Statistical Natural Language Processing.]]></article-title>

          <source>Foundations of Statistical Natural Language Processing.</source>

          <volume> </volume>

          <fpage></fpage>

          <lpage></lpage>

        </citation>

      </ref>


















      <ref id="102219">

        <label>6</label>

        <citation citation-type="book" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Van Rijsbergen, C.J.,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1989</year>

          <article-title><![CDATA[Information Retrieval.]]></article-title>

          <source>Information Retrieval.</source>

          <volume>2nd Edn.,</volume>

          <fpage>Pages: 323</fpage>

          <lpage>Pages: 323</lpage>

        </citation>

      </ref>






















      <ref id="4555">

        <label>7</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Agrawal, R. and R. Srikant,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1994</year>

          <article-title><![CDATA[Fast algorithms for mining association rules in large databases.]]></article-title>

          <source>Proceedings of the 20th International Conference on Very Large Data Bases,</source>

          <volume>1994</volume>

          <fpage>pp: 487</fpage>

          <lpage>499</lpage>

        </citation>

      </ref>


















      <ref id="50194">

        <label>8</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Cadez, I.V., S. Gaffney and P. Smyth,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2000</year>

          <article-title><![CDATA[A general probabilistic framework for clustering individuals and objects.]]></article-title>

          <source>Proceedings of the 6th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,</source>

          <volume>2000</volume>

          <fpage>pp: 140</fpage>

          <lpage>149</lpage>

        </citation>

      </ref>


















      <ref id="17699">

        <label>9</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Aggarwal, C.C., C.S. Gates and P.S. Yu,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1999</year>

          <article-title><![CDATA[On the merits of building categorization systems by supervised clustering.]]></article-title>

          <source>Proceedings of the 5th Conference on ACM Special Interest Group on Knowledge Discovery and Datamining,</source>

          <volume>1999</volume>

          <fpage>pp: 352</fpage>

          <lpage>356</lpage>

        </citation>

      </ref>


















      <ref id="50195">

        <label>10</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Gao, J., P.N. Tan and H. Cheng,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2006</year>

          <article-title><![CDATA[Semi-supervised clustering with partial background information.]]></article-title>

          <source>Proceedings of the 6th SIAM International Conference on Data Mining,</source>

          <volume>2006</volume>

          <fpage>pp: 487</fpage>

          <lpage>491</lpage>

        </citation>

      </ref>


















      <ref id="40594">

        <label>11</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Hotho, A., A. Maedche and S. Staab,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2001</year>

          <article-title><![CDATA[Text clustering based on good aggregations.]]></article-title>

          <source>Proceedings of the 2001 IEEE International Conference on Data Mining,</source>

          <volume>2001</volume>

          <fpage>pp: 607</fpage>

          <lpage>608</lpage>

        </citation>

      </ref>


















      <ref id="40419">

        <label>12</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Hu, X., X. Zhang, C. Lu and X. Zhou,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2009</year>

          <article-title><![CDATA[Exploiting wikipedia as external knowledge for document clustering.]]></article-title>

          <source>Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,</source>

          <volume>2009</volume>

          <fpage>pp: 389</fpage>

          <lpage>396</lpage>

        </citation>

      </ref>


















      <ref id="50196">

        <label>13</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Liu, L., E. Li, Y. Zhang and Z. Tang,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2007</year>

          <article-title><![CDATA[Optimization of frequent itemset mining on multiple-core processor.]]></article-title>

          <source>Proceedings of the 33rd International Conference on Very Large Data Bases,</source>

          <volume>2007</volume>

          <fpage>pp: 1275</fpage>

          <lpage>1285</lpage>

        </citation>

      </ref>


















      <ref id="50197">

        <label>14</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Pramudiono, I. and M. Kitsuregawa,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2003</year>

          <article-title><![CDATA[Parallel FP-growth on PC cluster.]]></article-title>

          <source>Proceedings of the 7th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining,</source>

          <volume>2003</volume>

          <fpage>pp: 467</fpage>

          <lpage>473</lpage>

        </citation>

      </ref>


















      <ref id="18445">

        <label>15</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Yang, Y.M. and J. Pedersen,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1997</year>

          <article-title><![CDATA[A comparative study on feature selection in text categorization.]]></article-title>

          <source>Proceedings of the 14th International Conference on Machine Learning,</source>

          <volume>1997</volume>

          <fpage>pp: 412</fpage>

          <lpage>420</lpage>

        </citation>

      </ref>


















      <ref id="24785">

        <label>16</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Hotho, A., S. Staab and G. Stumme,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2003</year>

          <article-title><![CDATA[Wordnet improves text document clustering.]]></article-title>

          <source>Proceedings of the SIGIR 2003 Semantic Web Workshop,</source>

          <volume>2003</volume>

          <fpage>pp: 541</fpage>

          <lpage>544</lpage>

        </citation>

      </ref>


















      <ref id="50198">

        <label>17</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Sedding, J. and D. Kazakov,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2004</year>

          <article-title><![CDATA[WordNet-based text document clustering.]]></article-title>

          <source>Proceedings of the 3rd Workshop on Robust Methods in Analysis of Natural Language Data,</source>

          <volume>2004</volume>

          <fpage>pp: 104</fpage>

          <lpage>113</lpage>

        </citation>

      </ref>


















      <ref id="50199">

        <label>18</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Wu, Z. and M. Palmer,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1994</year>

          <article-title><![CDATA[Verbs semantics and lexical selection.]]></article-title>

          <source>Proceedings of the 32nd Annual Meeting on Association for Computational Linguistics,</source>

          <volume>1994</volume>

          <fpage>pp: 133</fpage>

          <lpage>138</lpage>

        </citation>

      </ref>


















      <ref id="8064">

        <label>19</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Cutting, D.R., D.R. Karger, J.O. Pedersen and J.W. Tukey,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1992</year>

          <article-title><![CDATA[Scatter/gather: A cluster-based approach to browsing large document collections.]]></article-title>

          <source>Proceedings of the 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval,</source>

          <volume>1992</volume>

          <fpage>pp: 318</fpage>

          <lpage>329</lpage>

        </citation>

      </ref>


















      <ref id="49894">

        <label>20</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Zamir, O., O. Etzioni, O. Madani and R.M. Karp,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1997</year>

          <article-title><![CDATA[Fast and intuitive clustering of web documents.]]></article-title>

          <source>Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining,</source>

          <volume>1997</volume>

          <fpage>pp: 287</fpage>

          <lpage>290</lpage>

        </citation>

      </ref>


















      <ref id="50454">

        <label>21</label>

        <citation citation-type="conference" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Yang, X., D. Guo, X. Cao and J. Zhou,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>2008</year>

          <article-title><![CDATA[Research on ontology-based text clustering.]]></article-title>

          <source> Proceedings of the 2008 3rd International Workshop on Semantic Media Adaptation and Personalization,</source>

          <volume>2008</volume>

          <fpage>pp: 141</fpage>

          <lpage>146</lpage>

        </citation>

      </ref>












      <ref id="279774">

        <label>22</label>

        <citation citation-type="journal" xlink:type="simple">

          <person-group person-group-type="author">

            <name name-style="western">

              <surname>Gruber, T.R.,</surname>

              <given-names></given-names>

            </name>

          </person-group>

          <year>1993</year>

          <article-title><![CDATA[A translation approach to portable ontology specifications.]]></article-title>

          <source>Knowledge Acquisit.,</source>

          <volume>5</volume>

          <fpage>199</fpage>

          <lpage>220</lpage>

        </citation>

      </ref>
















  </ref-list>

</article>

