• [email protected]
  • +971 507 888 742
Submit Manuscript
SciAlert
  • Home
  • Journals
  • Information
    • For Authors
    • For Referees
    • For Librarian
    • For Societies
  • Contact
  1. Information Technology Journal
  2. Vol 12 (12), 2013
  3. 2447-2453
  • Issues
    Online First Current Issue All Issues
  • Information About
    Aims and Scope Editorial Board Guide to Authors Article Processing Charges
    Submit a Manuscript

Information Technology Journal

Year: 2013 | Volume: 12 | Issue: 12 | Page No.: 2447-2453
DOI: 10.3923/itj.2013.2447.2453

Facebook Twitter Reddit Linkedin E-mail
Google Scholar ASCI
Research Article

A New Hybrid Schemes Combining Ontology and Clustering for Text Documents

S.C. Punitha
Department of Computer Science and Engineering, Karunya University, Coimbatore, India

V. Thavavel
Department of Computer Application, Karunya University, Coimbatore, India

M. Punithavalli
Department of Computer Application, Sri Ramakrishna College of Engineering, Coimbatore, India

Data mining is a process of analyzing data from different perspectives and summarizing it into valuable information. It consist of two activities such as clustering and classification. It mainly works with numeric data, text data and the web data. Text-based algorithms have problems when dealing with different languages (synonyms, homonyms). Also, web pages contain other forms of information except text, such as images or multimedia. As a consequence, hybrid document clustering approaches have been proposed in order to combine the advantages and limit the disadvantages of the existing approaches. The main motivation behind ontology is that different people have different needs with regard to the clustering of texts. The hybrid schemes are developed using ontology and the frequent item clustering of various algorithms Ontology Based Apriori Based Clustering, Ontology based FP-Growth Based Clustering, Ontology based FP-Bonsai Clustering Algorithm have been proposed to resolve the disadvantages of existing approaches. The performance of this enhanced document clustering algorithm was tested vigorously using different datasets with performance measures to show the efficiency in clustering. Hence Ontology based FP-Bonsai Clustering Algorithm (OFPBC) shows significant improvement in terms of purity of clustering. The result shows that the datasets namely Reuters 21578,20 new Group and TDT2 which results the accuracy 0.840, 0.817 and 0.847 in OFPBC, respectively.
PDF Fulltext XML References Citation

How to cite this article

S.C. Punitha, V. Thavavel and M. Punithavalli, 2013. A New Hybrid Schemes Combining Ontology and Clustering for Text Documents. Information Technology Journal, 12: 2447-2453.

DOI: 10.3923/itj.2013.2447.2453

URL: https://scialert.net/abstract/?doi=itj.2013.2447.2453

Leave a Comment


Your email address will not be published. Required fields are marked *

Article Trend



Total views 1952

References


  1. Buckley, C. and A.F. Lewit, 1985. Optimization of inverted vector searches. Proceedings of the 8th Annual International SIGIR Conference on Research and Development in Information Retrieval, June 13-15, 1985, Montreal, Canada, pp: 97-110.

  2. Karypis, G., E.H. Han and V.K.P. Kumar, 1999. Chameleon: Hierarchical clustering using dynamic modeling. IEEE Comput., 32: 68-75.
    Direct Link

  3. Euzenat, J. and P. Shvaiko, 2007. Ontology Matching. 1st Edn., Springer-Verlag, Berlin, Germany, ISBN-13: 9783540496120, Pages: 333.

  4. Kowalski, G., 1997. Information Retrieval Systems: Theory and Implementation. 1st Edn., Kluwer Academic Publishers, Norwell, MA, USA., ISBN-13: 9780585320908, Pages: 282.

  5. Manning, C. and H. Schutze, 1999. Foundations of Statistical Natural Language Processing. MIT Press, Cambridge.

  6. Van Rijsbergen, C.J., 1989. Information Retrieval. 2nd Edn., Buttersworth Publishers, London, UK., Pages: 323.

  7. Agrawal, R. and R. Srikant, 1994. Fast algorithms for mining association rules in large databases. Proceedings of the 20th International Conference on Very Large Data Bases, September 12-15, 1994, San Francisco, USA., pp: 487-499.
    Direct Link

  8. Cadez, I.V., S. Gaffney and P. Smyth, 2000. A general probabilistic framework for clustering individuals and objects. Proceedings of the 6th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, August 20-23, 2000, Boston, MA, USA., pp: 140-149.

  9. Aggarwal, C.C., C.S. Gates and P.S. Yu, 1999. On the merits of building categorization systems by supervised clustering. Proceedings of the 5th Conference on ACM Special Interest Group on Knowledge Discovery and Datamining, August 15-18, 1999, San Diego, California, United States, pp: 352-356.
    Direct Link

  10. Gao, J., P.N. Tan and H. Cheng, 2006. Semi-supervised clustering with partial background information. Proceedings of the 6th SIAM International Conference on Data Mining, April 22, 2006, Bethesda, Maryland, pp: 487-491.

  11. Hotho, A., A. Maedche and S. Staab, 2001. Text clustering based on good aggregations. Proceedings of the 2001 IEEE International Conference on Data Mining, November 29-December 2, 2001, San Jose, pp: 607-608.
    CrossRef

  12. Hu, X., X. Zhang, C. Lu and X. Zhou, 2009. Exploiting wikipedia as external knowledge for document clustering. Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, June 28-July 1, 2009, Paris, France, pp: 389-396.

  13. Liu, L., E. Li, Y. Zhang and Z. Tang, 2007. Optimization of frequent itemset mining on multiple-core processor. Proceedings of the 33rd International Conference on Very Large Data Bases, September 23-27, 2007, Vienna, Austria, pp: 1275-1285.

  14. Pramudiono, I. and M. Kitsuregawa, 2003. Parallel FP-growth on PC cluster. Proceedings of the 7th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, April 30-May 2, 2003, Seoul, Korea, pp: 467-473.

  15. Yang, Y.M. and J. Pedersen, 1997. A comparative study on feature selection in text categorization. Proceedings of the 14th International Conference on Machine Learning, July 8-12, 1997, Nashville, TN., USA., pp: 412-420.
    Direct Link

  16. Hotho, A., S. Staab and G. Stumme, 2003. Wordnet improves text document clustering. Proceedings of the SIGIR 2003 Semantic Web Workshop, July 28-August 1, 2003, Toronto, Canada, pp: 541-544.

  17. Sedding, J. and D. Kazakov, 2004. WordNet-based text document clustering. Proceedings of the 3rd Workshop on Robust Methods in Analysis of Natural Language Data, August 29, 2004, Geneva, Switzerland, pp: 104-113.

  18. Wu, Z. and M. Palmer, 1994. Verbs semantics and lexical selection. Proceedings of the 32nd Annual Meeting on Association for Computational Linguistics, June 27-30, 1994, Las Cruces, New Mexico, USA., pp: 133-138.
    CrossRef

  19. Cutting, D.R., D.R. Karger, J.O. Pedersen and J.W. Tukey, 1992. Scatter/gather: A cluster-based approach to browsing large document collections. Proceedings of the 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, June 21-24, 1992, Copenhagen, Denmark, pp: 318-329.
    CrossRefDirect Link

  20. Zamir, O., O. Etzioni, O. Madani and R.M. Karp, 1997. Fast and intuitive clustering of web documents. Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining, August 14-17, 1997, Newport Beach, California, pp: 287-290.

  21. Yang, X., D. Guo, X. Cao and J. Zhou, 2008. Research on ontology-based text clustering. Proceedings of the 2008 3rd International Workshop on Semantic Media Adaptation and Personalization, December 15-16, 2008, IEEE Computer Society Washington, DC., USA., pp: 141-146.
    CrossRef

  22. Gruber, T.R., 1993. A translation approach to portable ontology specifications. Knowledge Acquisit., 5: 199-220.
    CrossRefDirect Link

Keywords


  • apriori algorithm
  • ontology
  • Document clustering
  • FP-growth algorithm

Useful Links

  • Journals
  • For Authors
  • For Referees
  • For Librarian
  • For Socities

Contact Us

Office Number 1128,
Tamani Arts Building,
Business Bay,
Deira, Dubai, UAE

Phone: +971 507 888 742
Email: [email protected]

About Science Alert

Science Alert is a technology platform and service provider for scholarly publishers, helping them to publish and distribute their content online. We provide a range of services, including hosting, design, and digital marketing, as well as analytics and other tools to help publishers understand their audience and optimize their content. Science Alert works with a wide variety of publishers, including academic societies, universities, and commercial publishers.

Follow Us
© Copyright Science Alert. All Rights Reserved