节点文献
基于向量空间模型的中文网页主题特征项抽取
Theme Feature Extraction of Chinese Webpage Based on Vector Space Model
【摘要】 为解决中文网页主题特征项抽取不精确的问题,对中文网页的主题特征项抽取算法进行了研究。网页的主题特征项抽取是主题网络爬虫进行网页相关度计算的基础,结合主题网页的二分类情况对目前常用的文本特征项加权方法 TF-IDF(Term Frequency-Inverse Document Frequency)进行了改进,在此基础上结合网页的半结构化特征,综合考虑特征项的位置信息及其包含的信息量,提出了一种线性特征项加权计算方法。经实验验证,该方法可有效提高主题网页的召回率和准确率。
【Abstract】 In order to solve the problem of imprecision in Chinese webpage theme feature extraction,feature extraction algorithm for Chinese webpage theme is studied. Webpage theme feature extraction is the foundation of topic web crawler to calculate webpage correlation. Considering two classifications of theme webpage,we improved the commonly used text feature item weighting method of TF-IDF( Term Frequency-Inverse Document Frequency). We combine Semi-structured characteristics of webpage,feature’s position information,present a new calculation method of linear feature item weighting. This method can effectively improve the theme webpage recall rate and precision rate.
【Key words】 term frequency-inverse document frequency(TF-IDF); vector space model; feature; correlation calculation; information gain;
- 【文献出处】 吉林大学学报(信息科学版) ,Journal of Jilin University(Information Science Edition) , 编辑部邮箱 ,2014年01期
- 【分类号】TP391.1
- 【被引频次】22
- 【下载频次】152