节点文献
Web文本内容过滤方法的研究
Research on Text Content Filtering in Web Pages
【Author】 YU Hai-yan, CHEN Xiao-Jiang, FENG Jian, FANG Ding-yi (Department of Computer Science, Northwest University, Xi’an 710069 China)
【机构】 西北大学计算机科学系;
【摘要】 文章研究了Web文本内容过滤的方法,分析了向量空间模型、关键词匹配算法等关键技术,并详细讨论了Web网页中文本内容过滤方法的实现过程。重点分析了该方法中的修正值选取、关键词权重函数以及过虑策略等方面的不足,提出了一个改进的Web文本内容过滤方法,能够有效降低算法的复杂性,提高性能。
【Abstract】 In this paper, the method of the text content filtering in Web pages is studied, the key techniques such as vector space model and keywords matching algorithm are analyzed, and then the implementing process of the text content filtering in Web pages is discussed in detail. The main point of this paper is the analysis of the shortages of this method about the revision value selection, the keywords weight function as well as the filtering of the strategy. An improved method of the text content filtering in Web pages is presented which can reduce the complexity of the algorithm efficiently and enhance it’s performance.
【Key words】 Text content filtering; Text vector; Keywords matching; Weight of keywords;
- 【会议录名称】 2006年全国开放式分布与并行计算学术会议论文集(一)
- 【会议名称】2006年全国开放式分布与并行计算学术会议
- 【会议时间】2006-10
- 【会议地点】中国陕西西安
- 【分类号】TP391.1
- 【主办单位】中国计算机学会开放系统专业委员会