节点文献
无双语词典的英汉词对齐
Aligning English-Chinese Words Without Bilingual Dictionary
【摘要】 该文提出了一种基于语料库的无双语词典的英汉词对齐模型 .它把自然语言的句子形式化地表示为集合 ,通过集合的交运算和差运算实现单词对齐 ,同时还考虑了词序和重复词的影响 .该模型不仅能对齐高频单词 ,而且能对齐低频单词 ,对未登录词和汉语分词错误具有兼容能力 .该模型几乎不需要任何语言学知识和语言学资源 ,使语料库方法可独立应用 .实验表明 ,同质语料规模越大 ,词对齐的正确率和召回率越高 .
【Abstract】 One of the bilingual corpus processing methods is the alignment of two languages on each linguistic level. Much research on word alignment between Indo-European languages has been done before, however, much less has been done on English-Chinese alignment. This paper proposes a corpus-based model for word alignment between English and Chinese. It formalizes natural languages into sets, and the intersection and difference of the sets to implement the word alignment. At the same time, the effect of word order and repetition is considered. The model includes a set of sub-models: minimum intersection model, minimum difference model, hybrid model, mono-directional model, bi-directional model, union model, and surrounding model. The English→Chinese mono-directional model is used to generate 1-m parallels, and the English←Chinese model is used to generate n-1 parallels. The union model and surrounding model are used to generate n-m parallels from the 1-m and n-1 parallels. The intersection of any two generated parallels in a sentence pair is empty, and the parallels themselves are minimum. This method can be used for alignment of both high-frequency words and low-frequency words, and is tolerant with Chinese word segmentation errors and unknown words. The typical characteristic of this model is that it needs few linguistic knowledge and resource. Experimental results show that the larger is the homogeneous corpus scale, the higher precision and recall rate can be obtained.
【Key words】 natural language processing; bilingual corpora; word alignment; minimum intersection; minimum difference;
- 【文献出处】 计算机学报 ,Chinese Journal of Computers , 编辑部邮箱 ,2004年08期
- 【分类号】TP391.1
- 【被引频次】33
- 【下载频次】413