节点文献

基于超链接信息的Web文本聚类方法研究

Research on the Method of Clustering Web Documents Based on Hyperlink Information

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙莉娜

【Author】 SUN Li-na (1.Tianjin University Electronic Information Engineering,Tianjin 300072,China;2.Shenyang School of Equipment Manufacturing Engineering,Liaoning 110026,China)

【机构】 天津大学电信学院 天津300072 沈阳市装备制造工程学校辽宁沈阳110026

【摘要】 面对当前大量的文本数据信息,如何帮助人们准确定位所需信息,成为文本挖掘领域的一个研究趋势。通过将文本分类和聚类方法应用于信息检索-—对网页文本进行聚类,提出了基于超链接信息的Web文本自动聚类模型。利用结构挖掘技术获得主题领域的多个权威网页作为初始聚类中心,通过去除超链接信息中的噪声和多余链接得到网站的简明拓扑结构,并结合内容挖掘,动态调整聚类中心,最终将网页聚成各主题下的不同子类别。

【Abstract】 Facing the massive volume text data information, how to locate the required information is one of the important research directions of text mining. The algorithms of text classification and clustering are applied to information retrieval, so the method of clustering Web documents based on hyperlink is presented according to the especial feature. And then the topological structure of website are found through hyperlink information, those noise and surplus hyperlink are cut down, the clusters are carried out based on the similarity between characteristic vectors which get from the content excavate of hyperlink anchor texts and web page texts. At the same time, the cluster centurions are adjusted dynamically, so as to realize the Web documents clustering based on hyperlink.

  • 【文献出处】 电脑知识与技术 ,Computer Knowledge and Technology , 编辑部邮箱 ,2006年26期
  • 【分类号】TP393.092
  • 【下载频次】143
节点文献中: 

本文链接的文献网络图示:

本文的引文网络