节点文献

基于支持向量机的音字转换模型

Pinyin-to-Character Conversion Model Based on Support Vector Machines

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 姜维关毅王晓龙刘秉权

【Author】 JIANG Wei,GUAN Yi,WANG Xiao-long,LIU Bing-quan (School of Computer Science and Technology,Haerbin Institute of Technology,Haerbin,Heilongjiang 150001,China)

【机构】 哈尔滨工业大学计算机科学与技术学院哈尔滨工业大学计算机科学与技术学院 黑龙江哈尔滨150001黑龙江哈尔滨150001

【摘要】 针对N-gram在音字转换中不易融合更多特征,本文提出了一种基于支持向量机(SVM)的音字转换模型,有效提供可以融合多种知识源的音字转换框架。同时,SVM优越的泛化能力减轻了传统模型易于过度拟合的问题,而通过软间隔分类又在一定程度上克服小样本中噪声问题。此外,本文利用粗糙集理论提取复杂特征以及长距离特征,并将其融合于SVM模型中,克服了传统模型难于实现远距离约束的问题。实验结果表明,基于SVM音字转换模型比传统采用绝对平滑算法的Trigram模型精度提高了1.2%;增加远距离特征的SVM模型精度提高1.6%。

【Abstract】 In order to overcome the difficulty in fusing more features into n-gram,a Pinyin-to-Character conversion model based on Support Vector Machines(SVM) is proposed in this paper,providing the ability of integrating more statistical information.Meanwhile,the excellent generalization performance effectively overcomes the overfitting problem existing in the traditional model,and the soft margin strategy overcomes the noise problem to some extent in the corpus.Furthermore,rough set theory is applied to extract complicated and long distance features,which are fused into SVM model as a new kind of feature,and solve the problem that traditional models suffer from fusing long distance dependency.The experimental result showed that this SVM Pinyin-to-Character conversion model achieved 1.2% higher precision than the trigram model,which adopted absolute smoothing algorithm,moreover,the SVM model with long distance features achieved 1.6% higher accuracy.

【基金】 国家自然科学基金重点项目资助(60435020);国家自然科学基金项目资助(60504021)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2007年02期
  • 【分类号】TP391.42
  • 【被引频次】17
  • 【下载频次】254
节点文献中: 

本文链接的文献网络图示:

本文的引文网络