节点文献

基于主题的Web文档聚类研究

Study on Topic-Based Web Clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙学刚陈群秀马亮

【Author】 SUN Xue-gang,CHEN Qun-xiu,MA Liang (State Key Laboratory of Intelligent Technology and System Dept. of Computer Science & Technology, Tsinghua University,Beijing 100084,China)

【机构】 智能技术与系统国家重点实验室清华大学计算机科学与技术系智能技术与系统国家重点实验室清华大学计算机科学与技术系 北京100084北京100084北京100084

【摘要】 网络资源的不断膨胀和新旧信息的迅速更迭 ,使传统的手工分检的方法难以适应对海量电子数据的管理需要。Web文档聚类可以快速地将文档进行自动归类 ,并能够发现新的信息资源。针对Web文档数据的复杂性 ,本文提出了通过二次特征提取和聚类的方法 ,将Web文档按照主题进行自动聚类。在主题特征被有效提取的同时 ,实现了较高质量的Web文档聚类。

【Abstract】 With the ceaseless resource inflation and rapid change of information on Web, it has become difficult to manage vast e-data through traditional manual method. Web clustering can automatically classify documents and help us to discover new information. Considering the complexity of Web documents, we offer a method of feature re-select and document re-cluster and perform a good Web clustering.

【基金】 国家 8 63资助项目 ( 2 0 0 1AA1140 4 0 )
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2003年03期
  • 【分类号】TP393.092
  • 【被引频次】82
  • 【下载频次】542
节点文献中: