节点文献
基于短语匹配的中文Web文档聚类算法
Phrase-Based Document Clustering Algorithm for Chinese Web Documents
【Author】 Wang Yang Zhang Lei Zhang Yi Computational Intelligence Laboratory, School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu 610054
【机构】 成都电子科技大学计算机科学与工程学院计算智能实验室;
【摘要】 本文在一种采用图结构、基于短语的文档索引模型的基础上,提出了一种基于短语匹配的、在线的、无需进行中文分词的增量聚类算法来对中文搜索结果进行聚类。结合文档索引模型和该聚类算法,可以有效地完成对搜索引擎所产生结果的增量式自动分类。
【Abstract】 An incremental clustering algorithm that is phrase-based, online and avoiding Chinese word segmentation has been proposed based on phrase-based document index model that is incremental constructed using graph technique for Chinese searching result. The auto-incremental classification of searching results can perform effectively and efficiently by combine the model and the algorithm.
【Key words】 Web Mining; Document Similarity; Phrase-based Matching; Document Clustering; Search Engine;
- 【会议录名称】 第二届全国信息检索与内容安全学术会议(NCIRCS-2005)论文集
- 【会议名称】第二届全国信息检索与内容安全学术会议(NCIRCS-2005)
- 【会议时间】2005-10
- 【会议地点】中国北京
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会信息检索与内容安全专业委员会