节点文献
论汉英平行语料的平行处理
Parallel Processing on Parallel Corpus of Chinese-English
【作者】 冯敏萱;
【导师】 陈小荷;
【作者基本信息】 南京师范大学 , 语言学及应用语言学, 2006, 博士
【摘要】 平行语料库研究是近年来语料库语言学横向发展的新趋势。人们已经清楚认识到大规模的高质量汉英平行语料库在自然语言处理、比较语言学研究和第二语言教学等众多领域中的巨大价值。但与单语语料库相比,汉英平行语料库无论在规模还是质量上都有较大差距。 为了进一步提高汉英平行语料的加工精度以适应建设和利用大规模平行语料的要求,本文以汉英平行语料的平行处理为主要研究对象,旨在利用双语信息,尤其是来自另一语言的信息来解决平行语料中某一语言的歧义问题。 本项研究主要取得了以下几方面成果: 第一,系统研究了平行处理技术。不仅明确了平行处理的含义,它在平行语料加工中的地位及价值,以及平行语料中用于消歧的语言资源层次及类别等等,而且还通过实验详细论证了平行处理技术在未登录词识别、词性标注、词义标注及句法分析等自然语言处理各层面的利用方法及有效性。 第二,平行处理技术是汉—英和英—汉双向的。我们不仅利用英语来解决汉语的歧义问题,包括汉语未登录词识别、汉语兼类词和多义词标注以及汉语“动词+名词”短语类型识别等,而且也利用汉语来解决英语歧义,例如英语的词性消歧和词义消歧等。 第三,在未经词汇对齐的平行语料中,实践了基于个性规则的词性、词义消歧方法。统计模型适于自动处理数据密集的问题,本文对英语人名汉译名的平行识别就主要使用了统计方法,精确率达到99.45%。而对于一些统计处理消歧效果较差、但出现频率又很高的词语,我们手工编写针对性极强的消歧规则。这些规则具有不受上下文长度和模板数量限制、特别适合于双语平行处理、消歧效果好等优点。我们为5个典型兼类词(过去、计划、与、back、so)和5个典型多义词(地方、所有、等、since、state)设计的平行处理算法,在大规模英汉或汉英平行语料中得到了验证,观察语料中的标注精确率均为100%,各类型语料中的总体精确率最高为100%,最低的也达到了96.59%,这比目前仅利用单语进行词性和词义消歧的成绩有了大幅度提高。 第四,精加工了1000句对的汉英平行语料。我们首先统计分析了这1000句对中汉英双语的词频、字词录入错误、普通未登录词、兼类词和多义词以及汉语的分词歧义字段、“动词+名词”序列等信息,然后利用平行处理技术,结合人工校对,消除了其中全部的句对齐、字词录入、分词和词性j际注错误,以此作为今后建设和加工大规模平行语料的可信资源。 综上所述,统计和规则相结合的平行处理技术,可以有效解决平行语料库中汉语或英语在单语处理时的许多困难问题,有利于更好地实现汉英机器翻译知识的自动获取。
【Abstract】 The research on parallel corpus is a new trend for corpus linguistics horizontal development. It has being known that the large Chinese-English parallel corpus of high quality has great value in the fields of the natural languages processing, the research of comparative linguistics and the teaching of second language, etc. But when compared with monolanguage corpus, the scale and quality of the Chinese-English parallel corpus are still far from users’ satisfaction.In order to improve the processing precision of the Chinese-English parallel corpus, so as to meet the requirement of the construction and using of large scale parallel corpus, this dissertation tries to take the parallel processing to Chinese-English parallel corpus as the main researching target and make use of bilingual information, especially the information from another language to solve ambiguities of one language among the parallel corpus.The following achievements have been obtained through this research:1. A systematic research of the parallel processing technique has been performed. The research has not only defined the meaning of the parallel processing, its position and value in the processing of the parallel corpus, the levels and types of language resources among the parallel corpus, which were used in disambiguation, but also demonstrated the applying approaches and validity of the parallel processing technique on each level of the natural language processing in detail, such as the recognition of the unknown words, the tagging of POS, the tagging of word sense and the syntactic analysis.2. The parallel processing technique is bidirectional (Chinese-English / English-Chinese). We have not only made use of English to settle the ambiguities in Chinese, including the recognition of Chinese unknown words, the tagging of Chinese words with polysemy and the recognition of phrasal type for Chinese phrases of "Verb+Noun", but also settled the ambiguities in English by using Chinese, such as the disambiguation of English POS and word sense.3. The POS and word sense disambiguation approaches based on individual rules have been experienced in non-lexical paralleled parallel corpus. The statistic models are suitable for processing the problem of concentrated data. This dissertation used statistic approach to carry out the parallel recognition of the Chinese translation of English