• [email protected]
  • +971 507 888 742
Submit Manuscript
SciAlert
  • Home
  • Journals
  • Information
    • For Authors
    • For Referees
    • For Librarian
    • For Societies
  • Contact
  1. Information Technology Journal
  2. Vol 8 (4), 2009
  3. 453-464
  • Issues
    Online First Current Issue All Issues
  • Information About
    Aims and Scope Editorial Board Guide to Authors Article Processing Charges
    Submit a Manuscript

Information Technology Journal

Year: 2009 | Volume: 8 | Issue: 4 | Page No.: 453-464
DOI: 10.3923/itj.2009.453.464

Facebook Twitter Reddit Linkedin E-mail
Google Scholar ASCI
Research Article

A Tolerance Rough Set Based Semantic Clustering Method for Web Search Results

Xian-Jun Meng
Intelligence Computing Research Center, Harbin Institute of Technology, Shenzhen Graduate School, 518055, People`s Republic of China

Qing-Cai Chen
Intelligence Computing Research Center, Harbin Institute of Technology, Shenzhen Graduate School, 518055, People`s Republic of China

Xiao-Long Wang
Intelligence Computing Research Center, Harbin Institute of Technology, Shenzhen Graduate School, 518055, People`s Republic of China

The objective of this study is to present a new web search results clustering algorithm which uses the tolerance rough set based approach to find the different meanings of the query in web search results and then organizes these results into different clusters according to their related meanings about query. Each meaning of the query can be represented by its contexts in each result and if there is a significant correlation between two context words, it is more likely that these two words represent the same meaning of query and also suitable as good indication of the meaning of query. In this study, the search results are organized in groups that each group of results relates to context words with high correlations and then these groups are merged into the final clusters representation using both cluster contents similarity and cluster documents overlap. The correlated context words with high documents coverage are selected as the labels of each cluster. Some experiments were conducted on different search results sets based on various queries. The results and comparisons of the proposed algorithm with that of the popular search results clustering algorithms through an empirical evaluation establish the viability of this proposed approach.
PDF Fulltext XML References Citation

How to cite this article

Xian-Jun Meng, Qing-Cai Chen and Xiao-Long Wang, 2009. A Tolerance Rough Set Based Semantic Clustering Method for Web Search Results. Information Technology Journal, 8: 453-464.

DOI: 10.3923/itj.2009.453.464

URL: https://scialert.net/abstract/?doi=itj.2009.453.464

Related Articles

A Rough Sets Based Data Preprocessing Algorithm for Web Structure Mining

Leave a Comment


Your email address will not be published. Required fields are marked *

Article Trend



Total views 3193

References


  1. An, A., Y. Huang, X. Huang and N. Cercone, 2004. Feature selection with rough sets for web page classification. Trans. Rough Sets, 2: 1-13.
    CrossRef

  2. Church, K. and P. Hanks, 1990. Word association norms, mutual information and lexicography. Comput. Linguist., 16: 22-29.
    CrossRef

  3. Cilibrasi, R.L. and P.M.B. Vitanyi, 2007. The Google similarity distance. IEEE Trans. Knowl. Data Eng., 19: 370-383.
    CrossRef

  4. Crabtree, D., P. Andreae and X. Gao, 2006. Query directed web page clustering. Proceedings of the 2006 IEEE/WIC/ACM International Conference on Web Intelligence, December 18-22, 2006, Hong Kong, pp: 202-210.
    CrossRef

  5. Dunning, T., 1993. Accurate methods for the statistics of surprise and coincidence. Comput. Linguist., 19: 61-74.
    Direct Link

  6. Francis, H., 2001. Mining associative meanings from the web: From word disambiguation to the global brain. Proceedings of the Trends in Special Language and Language Technology, March 29-30, 2001, Standaard Publishers Brussels, pp: 15-44.
    Direct Link

  7. Funakoshi, K. and T. Ho, 1998. A Rough Set Approach to Information Retrieval. In: Rough Sets in Knowledge Discovery, Polkowski, L. and A. Skowron (Eds.). Physica-Verlag, USA., ISBN: 978-3790811209, pp: 166-177.

  8. Hearst, M.A. and J.O. Pedersen, 1996. Reexamining the cluster hypothesis: Scatter/gather on retrieval results. Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, August 18-22, 1996, Zurich, Switzerland, pp: 76-84.
    CrossRefDirect Link

  9. Ho, T.B. and K. Funakoshi, 1998. Information retrieval using rough sets. J. Japan Soc. Artif. Intell., 13: 424-433.
    Direct Link

  10. Ho, T.B. and N.B. Nguyen, 2002. Nonhierarchical document clustering based on a tolerance rough set model. Int. J. Intell. Syst., 17: 199-212.
    CrossRef

  11. Jain, A.K., M.N. Murty and P.J. Flynn, 1999. Data clustering: A review. ACM Comput. Surv., 31: 264-323.
    CrossRefDirect Link

  12. Kawasaki, S., N.B. Nguyen and T.B. Ho, 2000. Hierarchical document clustering based on tolerance rough set model. Proceedings of the 4th European Conference on Principles of Data Mining and Knowledge Discovery, September 13-16, 2000, Lyon, France, pp: 458-463.
    CrossRef

  13. Kummamuru, K., R. Lotlikar, S. Roy, K. Singal and R. Krishnapuram, 2004. A hierarchical monothetic document clustering algorithm for summarization and browsing search results. Proceedings of the 13th International Conference on World Wide Web, May 17-20, 2004, ACM New York, USA., pp: 658-665.
    Direct Link

  14. Landauer, T.K. and S.T. Dumais, 1997. A solution to Plato's problem: The latent semantic analysis theory of acquisition, induction and representation of knowledge. Psychol. Rev., 104: 211-240.
    CrossRefDirect Link

  15. Leuski, A., 2001. Evaluating document clustering for interactive information retrieval. Proceedings of the 10th International Conference on Information and Knowledge Management, October 5-10, 2001, ACM Atlanta, Georgia, USA., pp: 33-40.
    CrossRef

  16. Lindsey, R., V. Veksler, A. Grintsvayg and W. Gray, 2007. Be wary of what your computer reads: the effects of corpus selection on measuring semantic relatedness. Proceedings of the 8th International Conference on Cognitive Modeling, July 27-29, 2007, Erlbaum Oxford, UK., pp: 279-284.
    Direct Link

  17. Lingras, P., 2002. Rough set clustering for web mining Fuzzy Syst., 2: 1039-1044.
    CrossRef

  18. Mecca, G., S. Raunich and A. Pappalardo, 2007. A new algorithm for clustering search results. Data Knowl. Eng., 62: 504-522.
    CrossRefDirect Link

  19. Miller, G., R. Beckwith, C. Fellbaum, D. Gross and K. Miller, 1990. Introduction to word Net: An on-line lexical database. Int. J. Lexicography, 3: 235-244.
    CrossRef

  20. Miller, G. and W. Charles, 1991. Contextual correlates of semantic similarity. Lang. Cogn. Proc., 6: 1-28.
    CrossRef

  21. Ngo, C.L. and H.S. Nguyen, 2005. A method of Web search result clustering based on rough sets. Proceedings of the 2005 IEEE/WIC/ACM International Conference on Web Intelligence, September 19-22, 2005, IEEE Computer Society, Washington, DC. USA., pp: 673-679.
    CrossRef

  22. Ohta, M., H. Narita and S. Ohno, 2004. Overlapping clustering method using local and global importance of feature terms at NTCIR-4 Web Task. Working Notes NTCIR, 4: 37-44.
    Direct Link

  23. Osinski, S., J. Stefanowski and D. Weiss, 2004. Lingo: Search results clustering algorithm based on singular value decomposition. Proceedings of the International Conference on Intelligent Information Systems (IIPWM), May 17-20, 2004, Zakopane, Poland, pp: 359-367.
    Direct Link

  24. Pantel, P. and D. Lin, 2002. Discovering word senses from text. Proceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, July 23-26, 2002, ACM Edmonton, Alberta, Canada, pp: 613-619.
    CrossRef

  25. Pawlak, Z., 1991. Rough Sets: Theoretical Aspects of Reasoning about Data. 1st Edn., Kluwer Academic Publishers, London, UK., ISBN-13: 9780792314721.

  26. Pedersen, T. and A. Kulkarni, 2007. Discovering identities in web contexts with unsupervised clustering. Proceedings of the IJCAI-2007 Workshop on Analytics for Noisy Unstructured Text Data, January 8, 2007, Springer, pp: 23-30.
    Direct Link

  27. Porter, M.F., 2006. An algorithm for suffix stripping. Program: Electron. Lib. Inform. Syst., 40: 211-218.
    CrossRefDirect Link

  28. Skowron, A. and J. Stepaniuk, 1996. Tolerance approximation spaces. Fundam. Infor., 27: 245-253.
    Direct Link

  29. Zamir, O. and O. Etzioni, 1998. Web document clustering: A feasibility demonstration. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, August 24-28, 1998, ACM Melbourne, Australia, pp: 46-54.
    CrossRef

  30. Zamir, O. and O. Etzioni, 1999. Grouper: A dynamic clustering interface to Web search results. Comput. Networks, 31: 1361-1374.
    CrossRef

  31. Zeng, H., Q. He, Z. Chen, W. Ma and J. Ma, 2004. Learning to cluster web search results. Proceeding of the 27th Annual International ACM SIGIR Conference on Research and Development in Informing Retrieval, July 25-29, 2004, Sheffield, South Yorkshire, UK., pp: 210-217.
    CrossRef

  32. Pawlak, Z., 1982. Rough sets. Int. J. Comput. Inform. Sci., 11: 341-356.
    CrossRefDirect Link

Keywords


  • rough set theory
  • web mining
  • semantic similarity
  • part-of-speech tagging
  • Search results clustering

Useful Links

  • Journals
  • For Authors
  • For Referees
  • For Librarian
  • For Socities

Contact Us

Office Number 1128,
Tamani Arts Building,
Business Bay,
Deira, Dubai, UAE

Phone: +971 507 888 742
Email: [email protected]

About Science Alert

Science Alert is a technology platform and service provider for scholarly publishers, helping them to publish and distribute their content online. We provide a range of services, including hosting, design, and digital marketing, as well as analytics and other tools to help publishers understand their audience and optimize their content. Science Alert works with a wide variety of publishers, including academic societies, universities, and commercial publishers.

Follow Us
© Copyright Science Alert. All Rights Reserved