节点文献

基于模糊近似度的Web文本过滤模型

The Feature Acquiring Algorithm on The Web Text

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【摘要】 <正> 从1991年诞生以来,WWW(World Wide Web)得到了迅猛的发展,它已经成为拥有约3亿用户、400万站点的巨大分布式信息空间、它包含了技术资料、商业信息、新闻报道、娱乐信息等多种类别和形式的信息,资源分布很分散,且没有统一的管理和结构。如何快速、准确地从浩瀚的信息资源中提取用户所需要的信息已经成为一个新的研究课题。WWW上最多的就是文本信息,因此Web信息处理的核心就是如何处理这些Web文档。数据挖掘和知识发现(Data Mining and Knowl-edge Discovery,DMKD)可以帮助人们从大量原始数据中挖掘出隐含的、有用的尚未发现的信息和知识,有效地解决信息丰富知识贫乏问题。因此,基于Web文本信息的挖掘作为数据挖掘的一个新主题,引起了人们的极大兴趣。Web文本信息的挖掘就是在大量训练样本的基础上,得到文本数据间的内在特征,并以此为依据在网络资源中进行有目的的信息提取。在本文中,我们首先介绍了Web文本信息的向量空间表示模型(VSM),并在此模型的基础上提出了一

【Abstract】 The booming growth of the Internet provides us a great deal of information resource. In this paper, we create a text filtering model based on VSM. In this model, Web text mining is an efficient technique,which discoveres valuable and potential knowledge from those unstructured texts. In this paper, we use VSM as the description of Web text and give a feature subset algorithm which is based on the Genetic Algorthm. This algorithm can greatly improve the efficiency of dealing with Web texts and give much better way to classify and cluster the texts. Our experiments show that this method is active well in feature dimension reduction.

【关键词】 VSMText filteringGenetic algorithmText miningKDD
【Key words】 VSMText filteringGenetic algorithmText miningKDD
【基金】 天津自然科学基金(003700111)和(993600811)
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2001年12期
  • 【分类号】TP393.09
  • 【被引频次】6
  • 【下载频次】84
节点文献中: 

本文链接的文献网络图示:

本文的引文网络