节点文献

短文本数据的自动分类

Short-Text Categorization

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 宋东风张志浩

【Author】 SONG Dong-feng,ZHANG Zhi-hao(Department of Computer Science and Technology,Tongji University,Shanghai 200092,China)

【机构】 同济大学计算中心同济大学计算中心 上海200092上海200092

【摘要】 文章以比较购物搜索中的商品数据自动分类为应用背景,探讨短文本数据的分类问题,比较了常用的文本分类算法的特点,在此基础上提出k-NN与NB相结合的多分类器方案,对于NB算法分类不可信的情况下改用k-NN算法进行再次分类,并充分利用NB的中间结果供k-NN剪枝时作参考。实验数据表明该方法在与NB相近的时间复杂度下可明显地提高短文本分类的正确率和召回率,达到实际应用的要求。

【Abstract】 On the basis of the application of the automatism in the comparison shopping,this paper has probe into the issue of text categorization.It has compared two popular algorithms for text categorization:Naive Bayes(NB) and k-Nearest Neighbor(k-NN).On this basis it proposed another suggestion-combining the two algorithms.In the situation that NB is unauthentic,K-NN arithmetic is suggested to be used to recategorize the result.And the k-NN algorithm can also make best use of the results from the NB algorithm during the process of re-categorization.The statistics from the experiment show that under the similar time complexity,the new algorithm can improve markedly the precision of the text categorization and the recall rate.It can reach the expected demand.

【关键词】 文本分类短文本朴素贝页斯k近邻
【Key words】 text categorizationshort textNaive Bayesk-NN
  • 【文献出处】 电脑与信息技术 ,Computer and Information Technology , 编辑部邮箱 ,2007年01期
  • 【分类号】TP391.1
  • 【被引频次】11
  • 【下载频次】343
节点文献中: