节点文献

Web文本信息的特征获取算法

Feature Acquiring Algorithm on the Web Text

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘明吉王秀峰饶一梅黄亚楼

【Author】 LIU Ming ji, WANG Xiu feng, RAO Yi mei, HUANG Ya lou (Department of Computer and System Sciences, Nankai University, Tianjing 300071,China)

【机构】 南开大学计算机与系统科学系南开大学计算机与系统科学系 天津300071天津300071天津300071

【摘要】 Internet的发展为人们提供了大量的信息资源 ,Web文本挖掘是从非结构化的文本中发现潜在的、有价值知识的一种有效技术 .本文以矢量空间模型为 Web文本的表示方法 ,提出了一个基于遗传算法的 Web文本特征抽取算法 ,进一步提高了 Web文本的处理效率 ,为文本的分类、聚类以及其它处理提供了简练的特征表示方法 .实验证明 ,该种处理方法有效地降低了文本特征矢量的维数 .

【Abstract】 The booming growth of the Internet provides us a great deal of information resource. Web text mining is an efficient technique, which discovery valuable and potential knowledge from those unstructured texts. In this paper, we use VSM as the description of web text and give a feature subset algorithm which is based on the Genetic Algorithm. This algorithm can greatly improve the efficiency of dealing with web texts and give much better way to classify and cluster the texts. Our experiments show that this method active well in feature dimension reduction.

【关键词】 Web挖掘VSM遗传算法文本特征抽取
【Key words】 Web miningVSMgenetic algorithmtext feature abstract
【基金】 天津自然科学技术基金项目 (0 0 3 70 0 111)、(993 60 0 811)和 (0 0 3 60 0 3 11)资助
  • 【文献出处】 小型微型计算机系统 ,Mini-micro Systems , 编辑部邮箱 ,2002年06期
  • 【分类号】TP393.03
  • 【被引频次】95
  • 【下载频次】539
节点文献中: 

本文链接的文献网络图示:

本文的引文网络