节点文献

一种面向海量实时数据的信息检索算法

An Information Retrieval Algorithm for Massive and Real-time Data

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 丁伟林容容倪良胜

【Author】 Ding Wei Lin Rong-rong Ni Liang-sheng (Dept. of Computer Science & Engineering, Southeast Univ. , Nanjing 210096, Jiangsu, China)

【机构】 东南大学计算机科学与工程系东南大学计算机科学与工程系 江苏 南京 210096江苏 南京 210096江苏 南京 210096

【摘要】 网络信息资源的迅猛膨胀推进了信息检索技术的发展和成熟,但将现有的技术应用于海量实时网络数据时,传统的信息检索算法仍存在种种不足之处.本文中以CER-NET华(东)北地区的海量实时网络数据环境为依托,研究和设计了两段向量簇聚类信息检索算法,通过插入聚类和优化聚类两阶段的操作,提供高效的信息处理能力.同时,基于簇聚类树实现了群发邮件甄别的应用,对网络数据中的垃圾邮件进行过滤,进一步地提高检索效率.

【Abstract】 With the rapid expansion of information resources in networks, information retrieval technologies are now becoming more and more well-developed. But their current applications to massive and real-time data, especially for the conventional information retrieval algorithms, still reveal some shortcoming. In this paper, aiming at the massive and real-time network data from CERNET East China North center, a two-phase vector clustering algorithm is investigated and designed, in which a high-efficiency information processing ability is implemented by a two-phase operation; clustering insertion and clustering optimization. Meanwhile, the application of the proposed algorithm in the group mail discrimination system for filtering junk mails of network data is achieved by means of the clustering tree. Thus, the retrieval efficiency is further improved.

  • 【文献出处】 华南理工大学学报(自然科学版) ,Journal of South China University of Technology(Natural Science) , 编辑部邮箱 ,2004年S1期
  • 【分类号】TP393.09
  • 【被引频次】1
  • 【下载频次】281
节点文献中: 

本文链接的文献网络图示:

本文的引文网络