节点文献

统计和词典方法相结合的双语语料库词对齐

Word Alignment Based on Statistic and Lexicon

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吕雅娟赵铁军李生杨沐昀

【Author】 Lu Yajuan Zhao TieJun Li Sheng Yang MuYunDept. of Computer Science & Engineering, Harbin Institute of Technology, Harbin, 15001

【机构】 哈尔滨工业大学计算机科学与技术学院

【摘要】 双语语料库词对齐研究对于自然语言处理的许多应用具有重要意义.本文在对基于统计和基于词典的词对齐方法进行实验分析的基础上,提出了基于词典的语言学信息与词性统计相结合的词对齐方法,该方法在训练语料库规模较小的情况下,充分利用现有资源,取得了较好的词对齐结果.

【Abstract】 Word alignment of bilingual corpus are useful for many NLP applications. On comparing the word alignment results abtained from statistic-based and lexicon-based methods, this paper proposed a hybrid method that combine linguistic information and POS-based statistic. The proposed method achieve good performance based on a small training corpus.

【关键词】 自然语言处理双语语料库词对齐
【Key words】 NLPBilingual CorpusWord Alignment
【基金】 国家自然科学基金的资助,合同号69775017
  • 【会议录名称】 自然语言理解与机器翻译——全国第六届计算语言学联合学术会议论文集
  • 【会议名称】全国第六届计算语言学联合学术会议
  • 【会议时间】2001-08
  • 【会议地点】中国山西
  • 【分类号】H085
  • 【主办单位】山西大学计算机系
节点文献中: