节点文献

基于LSA和SVM的文本分类模型的研究

Research of text categorization model based on LSA and SVM

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王永智滕至阳王鹏聂江涛

【Author】 WANG Yong-zhi, TENG Zhi-yang, WANG Peng, NIE Jiang-tao (School of Computer Science and Engineering, Southeast University, Nanjing 210096, China)

【机构】 东南大学计算机科学与工程学院

【摘要】 为了提高文本分类的准确性,研究并设计了一个基于潜在语义分析和支持向量机的多类文本分类模型。利用潜在语义分析进行特征抽取,消除多义词和同义词在文本表示时造成的偏差,并实现文本向量的降维。使用具有良好分类精度和泛化能力的支持向量机进行分类,提出一种改进的一对一多类分类算法,改善不可分问题。实验结果表明,该模型在类别数目较少时具有较好的分类效果。

【Abstract】 A multiclass text categorization model based on latent semantic analysis and support vector machine is researched and designed to enhance the accuracy of categorization. Using latent semantic analysis to extract feature, the affect of synonymy and polysemy in text representation process is eliminated and the dimension of text vector is reduced. Support vector machine, which has high precision of categorization and excellent generalization ability, is used to categorize text. And an improved 1-v-1 method is presented to resolve the unclassifiable problem. Experimental results show that the proposed model has better categorization performance when the category number is small.

  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2009年03期
  • 【分类号】TP391.1
  • 【被引频次】24
  • 【下载频次】444
节点文献中: