节点文献

现代汉语动态助词“了”的自动生成研究

【作者】 何晓丽

【导师】 陈小荷;

【作者基本信息】 南京师范大学 , 语言学及应用语言学, 2007, 硕士

【摘要】 自然语言生成是当前以计算语言学和人工智能为基础的自然语言处理中相当活跃的一个分支,主要研究如何用计算机来生成自然语言文本,有着极其重要的应用价值。自然语言生成的研究可以作为检验特定语言理论的一种技术手段,不断为理论语言学提供反馈,推动语言学朝纵深方向发展。动态助词“了”一直是传统语言学界研究的热点和难点,而其特点和使用规则是难点之一,传统的研究方法多为语义和语法描写法,多是定性分析,且缺乏对动态助词“了”的全面考察,并仍存在不少分歧,因而缺乏整体的说服力,有的还需要重新考虑。而且,仍有许多问题没有解决,如现代汉语里动词带动态助词“了”的实际情况如何?有无规律性?如有规律,有何规律?本文主要探讨如何用自然语言生成方法自动在汉语的动词之后生成动态助词“了”。在传统语言学界已有研究成果的基础上,通过学习和借鉴前人的理论、方法、经验和教训,结合大规模真实语料,利用以规则为主,统计为辅,两者相结合的技术,从自然语言生成角度来考虑动态助词“了”的使用。语料的观察和统计是进行动态助词“了”生成研究的出发点。结合大规模语料库,对相当数量的动词带动态助词“了”的情况进行大量的考察,以此来统计现代汉语里动词带动态助词“了”的实际情况,来归纳总结动词带动态助词“了”的规律性等等。基于规则的生成策略是生成试验采用的的主要技术。在总结传统语言学的已有研究成果的基础上,通过对语料库的统计和观察来增进知识,完善生成规则库,形成两个主要的生成规则库:“不可加‘了/u’的规则”库和“可加‘了/u’的规则”库。在基于规则的生成系统中,通过对规则库中的规则进行有序组织来有效地解决规则间的冲突问题。按照不同的层次存放规则和尽可能细分每一类型的规则是主要策略。文中分别以标注为V的动词和句子为单位来衡量生成结果,数据更加有层次性,更加客观可信。同时,还考虑到汉语动词中复杂的可加可不加动态助词“了”的情形,采用两个底本来衡量生成结果。一是完全忠实于原文的硬性底本;二是加入了人工干预的弹性底本,有效提高了正确率。就精确率而言,封闭和开放测试都取得了较好的效果。

【Abstract】 Natural Language Generation is a quite active field of the natural language processing which is based on computational linguistics and artificial intelligence now.It studies how to use computer to generate natural language text, and it has extremely important using value. The research can be regarded as a kind of technological means to test the particular language theory, and offer feedbacking for theoretical linguistics constantly, promote linguistics to develop in the depth direction.Dynamic auxiliary word "le" has been the focus and difficult point of the traditional language study.The characteristic and the rule of using of the dynamic auxiliary word "le" is one of the difficult point.Though the linguists have already paid close attention to it, but its research approaches are almost traditional semanteme and grammar describing, are mostly qualitative analysis, and lack of an overall investigations,and it still exists much differences, therefore,it lacks of the convincingness of the whole, and some need to reconsider.Moreover, it still has a lot of problem to solve,for example,how about the actual conditions of the dynamic auxiliary word "le" in modern Chinese? Is there regularity? If it is regular, what laws are there?This text mainly discussed how to generate the dynamic auxiliary word "le" after the verbs automaticly by using the method of the Natural Language Generation. This text is on the basis of utilizing the achievements of the traditional language reseach now, through studying and using the theory, method, experience and lesson of forefathers for reference, combine the extensive true language material, utilize the technology which takes rule as the core, and uses statistics in order to complement and combine the two together, to reconsider the using of the dynamic auxiliary word "le" in the angle of natural language generation.Observation and statistics of the language material is the starting point of generating the dynamic auxiliary word "le". We investigate a considerable amount of "le" with verbs in extensive corpus,in order to statistic the actual situation of the using of the dynamic auxiliary word "le" after verbs,and to summarize the regularity of the verbs which can add "le",etc. The rules are formalized and optimized from observing and statisticing the corpus and summarizing the forefathers’ research results.Our main technology of the generating experiment is adopted on the basis of the regular generation strategy. On the basis of summarizing the research results of the traditional linguistics and promoting knowledge by statistics and observation of the corpus, we improve the create-rule storehouse, and form two main create-rules collection: "the rules which can’t add ’le’ " and " the rules which can add ’le’". In the rule-based generation system, we solve the problem of the conflict among the rules effectively through organizing the rule in order.To preserve the rule according to different levels and subdivide every type of the rules are the main tactics.We separately weigh the result of generating in two units,one is the verb which are marked "V " and the other is the big sentence,so the data have more levels, and they are more objective. Meanwhile,we also consider the complicated sitiatio n that it is proper for some verbs whether they add dynamic auxiliary words"le" or not. We adopt two bottom pieces to weigh the results of generating. The fist is a rigid copy which is totally faithful to the original text; The second is an elastic copywhich is joined manual intervention, but it has improved the correct rate greatly. Regards to the accurate rate, close test and open test both have good results.

  • 【分类号】H146.2
  • 【被引频次】6
  • 【下载频次】831
节点文献中: