节点文献
Web文本信息的特征获取算法
Feature Acquiring Algorithm on the Web Text
【摘要】 Internet的发展为人们提供了大量的信息资源 ,Web文本挖掘是从非结构化的文本中发现潜在的、有价值知识的一种有效技术 .本文以矢量空间模型为 Web文本的表示方法 ,提出了一个基于遗传算法的 Web文本特征抽取算法 ,进一步提高了 Web文本的处理效率 ,为文本的分类、聚类以及其它处理提供了简练的特征表示方法 .实验证明 ,该种处理方法有效地降低了文本特征矢量的维数 .
【Abstract】 The booming growth of the Internet provides us a great deal of information resource. Web text mining is an efficient technique, which discovery valuable and potential knowledge from those unstructured texts. In this paper, we use VSM as the description of web text and give a feature subset algorithm which is based on the Genetic Algorithm. This algorithm can greatly improve the efficiency of dealing with web texts and give much better way to classify and cluster the texts. Our experiments show that this method active well in feature dimension reduction.
【Key words】 Web mining; VSM; genetic algorithm; text feature abstract;
- 【文献出处】 小型微型计算机系统 ,Mini-micro Systems , 编辑部邮箱 ,2002年06期
- 【分类号】TP393.03
- 【被引频次】95
- 【下载频次】539