节点文献

平行语料库中双语术语词典的自动抽取

Automatic Extraction of Bilingual Term Lexicon from Parallel Corpora

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙乐金友兵杜林孙玉芳

【Author】 SUN Le JIN You bing DU Lin SUN Yu fang (Chinese Information Processing Center,Institute of Software,Chinese Academy of Sciences Beijing 100080) E mail:lesun,ybjin,ldu,yfsun@sonata.iscas.ac.cn

【机构】 中国科学院软件研究所中文信息处理中心!北京100080

【摘要】 本文提出了一种从英汉平行语料库中自动抽取术语词典的算法。首先采用基于字符长度的改进的统计方法对平行语料进行句子级的对齐 ,并对英文语料和中文语料分别进行词性标注和切分与词性标注。统计已对齐和标注的双语语料中的名词和名词短语生成候选术语集。然后对每个英文候选术语计算与其相关的中文翻译之间的翻译概率。最后通过设定随词频变化的阈值来选取中文翻译。在对真实语料的术语抽取实验中取得了较好的结果

【Abstract】 An algorithm for the automatic extraction of a bilingual term lexicon from English Chinese parallel corpora is proposed in this paper.Parallel corpora are firstly aligned by improved statistical method,which is based on character length,and tagged with their part of speech categories respectively.The term candidate set is produced by statistical the nouns and noun phrases of both corpora.Then the translation probability between every English candidate term and its Chinese translation term are calculated.Finally,the Chinese translation of English term is selected by threshold value,which varies with word frequency.A better performance is obtained in the experiments of term extraction on real corpora.

【基金】 国家青年自然科学基金!(6 99830 0 9)
  • 【文献出处】 中文信息学报 ,JOURNAL OF CHINESE INFORMATION PROCESSING , 编辑部邮箱 ,2000年06期
  • 【分类号】TP391
  • 【被引频次】78
  • 【下载频次】927
节点文献中: 

本文链接的文献网络图示:

本文的引文网络