节点文献

SVM词库智能更新技术在搜索分类中的应用

Research of thesaurus intelligent update technology based on SVM applied in search classification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 齐富民谢晓尧景凤宣

【Author】 QI Fu-min;XIE Xiao-yao;JING Feng-xuan;Key Laboratory of Information and Computing Science Guizhou Province,Guizhou Normal University;

【机构】 贵州师范大学贵州省信息与计算科学重点实验室

【摘要】 为了研究搜索引擎的文本预分类准确率,从词库对搜索引擎的影响角度出发,提出了基于支持向量机的词库智能更新技术。利用网络爬虫丰富的数据源作为生僻词来源,用基础词库结合语法库对网络爬虫获取的文本语料进行分析处理,同时不断充实临时词库;利用支持向量机判定文本的所属类别,确定生僻词的类别标识;根据临时词库中的生僻词的统计数量,将生僻词加入到词库,达到扩大词库的目的。将扩展后的词库应用于搜索引擎的搜索意图识别实验中,实验结果表明,扩展后的词库可以减少句子拆分的错误率并提高搜索主题分类的准确率。

【Abstract】 To research the accuracy of the texts’ pre-classification on search engine,the thesaurus intelligent update technology on support vector machine was proposed,and the rich data sources from the Web crawler as the data source of uncommon words was taken,and basic vocabulary library combining with the syntax library was used for processing and analyzing the text corpus grabbing from the Web crawler,meanwhile,enriching temporary thesaurus. SVM was used to distinguish the category of text and to determine the category identification of uncommon words. According to the statistical number of uncommon words in the provisional lexicon,the uncommon words was added to the lexicon,and the goal of expanding the thesaurus was reached. Practice proved that the technology used in the search engine text word and subject classification could reduce the error rate of sentence splitting and improve the accuracy rate of searching subject classification.

【基金】 贵州省工业攻关基金项目(黔科合GY字[2008]3009);贵州省科学技术基金项目(黔科合J字[2011]2213);贵州师范大学2012年度自然科学类学生科研基金重点项目(201219)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2014年06期
  • 【分类号】TP391.1
  • 【被引频次】2
  • 【下载频次】151
节点文献中: 

本文链接的文献网络图示:

本文的引文网络