节点文献

基于双语平行语料的中文缩略语提取方法

A Bilingual-constrained Approach for Extracting Chinese Abbreviations

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘友强李斌奚宁陈家骏

【Author】 Liu You-qiang~1,Li Bin~(1,2),Xi Ning~1,Chen Jia-jun~1 1 State Key Laboratory for Novel Software Technology at Nanjing University,Nanjing 210093 2Research Center of Language and Informatics,Nanjing Normal University,Nanjing 210097

【机构】 南京大学计算机软件新技术国家重点实验室南京师范大学语言信息科技研究中心

【摘要】 汉语缩略语在现代汉语中被广泛使用,其相关研究对于中文信息处理有着重要的意义。本文提出了一种从英汉平行语料库中自动提取汉语缩略语的方法。我们首先对双语语料进行词对齐训练,利用训练得到的词对齐信息抽取出候选中英文短语对。然后用SVM分类器提取出质量高的短语对。最后再从质量高的短语对集合中利用英文翻译及一些汉语缩略-全称对应规则提取出汉语缩略语及全称语对。实验结果表明,该方法提取出的缩略语具有较高的准确率,可以作为一种自动提取缩略语词典的有效方法。

【Abstract】 Chinese abbreviations are widely used in modem Chinese texts,and the correlated research is important for Chinese information processing.In this paper,we propose an approach to extract Chinese abbreviations from Chinese-English parallel corpus.First we generate word alignments for the corpus,and extract Chinese-English phrase pairs consistent with the alignments.After that,we discriminate high quality phrase pairs from the bad ones by SVM Classifier.Then we extract Chinese abbreviation and full-form phrase pairs from the high quality group using their corresponding English translations and some rules.The experiments showed that our approach can extract abbreviations with high accuracy,and could be an effective way to extract Chinese abbreviation and full-form phrase pairs.

【关键词】 缩略语平行语料库短语抽取分类
【Key words】 abbreviationparallel corpusphrase extractionclassify
【基金】 国家自然科学基金(61003112,61073119);国家社会科学基金(10CYY021);南京大学计算机软件新技术国家重点实验室(KFKT2011B03)的资助
  • 【会议录名称】 中国计算语言学研究前沿进展(2009-2011)
  • 【会议名称】第十一届全国计算语言学学术会议
  • 【会议时间】2011-08-20
  • 【会议地点】中国河南洛阳
  • 【分类号】TP391.1;H136.6
  • 【主办单位】中国中文信息学会
节点文献中: