节点文献

基于前后文n-gram模型的古汉语句子切分

Archaic Chinese Punctuating Sentences Based on Context N-gram Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈天莹陈蓉潘璐璐李红军于中华

【Author】 CHEN Tianying 1,CHEN Rong1,PAN Lulu1,LI Hongjun1,2,YU Zhonghua1(1.Dept.of Computer Science,Sichuan University,Chengdu 610064;2.Dept.of Computer Science,Southwest University of Science and Technology,Mianyang 621002)

【机构】 四川大学计算机学院四川大学计算机学院 成都610064成都610064西南科技大学计算机学院绵阳621002

【摘要】 提出了基于前后文n-gram模型的古汉语句子切分算法,该算法能够在数据稀疏的情况下,通过收集上下文信息,对切分位置进行比较准确的预测,从而较好地处理小规模训练语料的情况,降低数据稀疏对切分准确率的影响。采用《论语》对所提出的算法进行了句子切分实验,达到了81%的召回率和52%的准确率。

【Abstract】 An algorithm of punctuating the sentences in archaic Chinese language based on context n-gram model is proposed in the paper.The algorithm can make comparatively accurate prediction of the punctuating-positions of the text under data-sparse instances by collecting and calculating context information to better analyze small-scaled corpus and meanwhile,to bring down the effects of the data-sparse plight on the global accuracy.At last,the paper selects the analects of Confucius(Lunyu) to test the algorithm introduced,and the results show that the recall and the precision achieve 81% and 52% respectively.

【基金】 国家自然科学基金资助项目(60073046);高等学校博士学科点专项科研基金“SRFDP”资助项目(20020610007)
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2007年03期
  • 【分类号】TP391.1
  • 【被引频次】27
  • 【下载频次】352
节点文献中: