节点文献

基于SVM预测的金融主题爬虫

Financial topical crawler based on SVM prediction

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈黎李志蜀琚生根唐小棚梁时木韩国辉

【Author】 CHEN Li,LI Zhi-Shu,JU Sheng-Gen,TANG Xiao-Peng,LIANG Shi-Mu,HAN Guo-Hui (College of Computer Science,Sichuan University,Chengdu 610064,China)

【机构】 四川大学计算机学院

【摘要】 随着Internet上信息的爆炸,利用通用搜索引擎检索用户相关的信息变得越来越困难,而主题爬虫成为WEB上检索主题相关信息的重要工具.目前大部分基于分类器预测的主题爬虫的训练数据是不同类别网页的内容,但是在实际预测过程只能根据父网页中的一些链接信息进行预测,所以造成主题爬虫的预测的准确率较低.本文使用SVM分类器对标注了类别的URL以及上下文和锚文本进行训练,并分别使用了DF和信息增益两种不同的特征选择方法进行特征筛选,对影响分类器的各种因素进行了实验对比,并对分类器进行了在线的实验.实验证明这种方法在实际预测过程中效率很高.

【Abstract】 With the rapid growth of information and the explosion of web pages from the World Wide Web,it gets harder for general crawlers to retrieve the information relevant to a user.Topical crawlers are becoming important tools to gather web pages on a specific topic.Training set of topical crawler based on classifier prediction comes from different kinds of Web contents,but most of classifier can predict according to some links information of parent Web pages in actual condition.As being different kinds of information between training and testing,the accuracy of this kind of classifier is low.SVM classifier is used in this paper to train the contexts and anchors of URLs,and train different information from different character selection methods,the DF and information gain to contrast experiment results based on all sorts of factors which will impact on classifier.It can validate that there is of very high accuracy in actual prediction when classifier being on-line experiments.

【基金】 四川省科技厅公益性研究计划项目(2008SZ0049)
  • 【文献出处】 四川大学学报(自然科学版) ,Journal of Sichuan University(Natural Science Edition) , 编辑部邮箱 ,2010年03期
  • 【分类号】TP391.3
  • 【被引频次】7
  • 【下载频次】100
节点文献中: 

本文链接的文献网络图示:

本文的引文网络