节点文献

基于向量空间模型的中文网页主题特征项抽取

Theme Feature Extraction of Chinese Webpage Based on Vector Space Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 代宽赵辉韩冬宋天勇

【Author】 DAI Kuan;ZHAO Hui;HAN Dong;SONG Tian-yong;College of Computer Science and Engineering, Changchun University of Technology;College of Software Vocational Technology, Changchun University of Technology;

【机构】 长春工业大学计算机科学与工程学院长春工业大学软件职业技术学院

【摘要】 为解决中文网页主题特征项抽取不精确的问题,对中文网页的主题特征项抽取算法进行了研究。网页的主题特征项抽取是主题网络爬虫进行网页相关度计算的基础,结合主题网页的二分类情况对目前常用的文本特征项加权方法 TF-IDF(Term Frequency-Inverse Document Frequency)进行了改进,在此基础上结合网页的半结构化特征,综合考虑特征项的位置信息及其包含的信息量,提出了一种线性特征项加权计算方法。经实验验证,该方法可有效提高主题网页的召回率和准确率。

【Abstract】 In order to solve the problem of imprecision in Chinese webpage theme feature extraction,feature extraction algorithm for Chinese webpage theme is studied. Webpage theme feature extraction is the foundation of topic web crawler to calculate webpage correlation. Considering two classifications of theme webpage,we improved the commonly used text feature item weighting method of TF-IDF( Term Frequency-Inverse Document Frequency). We combine Semi-structured characteristics of webpage,feature’s position information,present a new calculation method of linear feature item weighting. This method can effectively improve the theme webpage recall rate and precision rate.

【基金】 吉林省科技厅自然科学基金资助项目(20130101060JC)
  • 【文献出处】 吉林大学学报(信息科学版) ,Journal of Jilin University(Information Science Edition) , 编辑部邮箱 ,2014年01期
  • 【分类号】TP391.1
  • 【被引频次】22
  • 【下载频次】152
节点文献中: 

本文链接的文献网络图示:

本文的引文网络