节点文献

语料库研究

Study on Corpus

【作者】 何婷婷

【导师】 李宇明;

【作者基本信息】 华中师范大学 , 语言学及应用语言学, 2003, 博士

【摘要】 本文以语料库本身为研究对象,以语言学理论为基础,以计算机软件工程和数据库的思想为指导,结合其他学科领域的理论和方法,在总结前人提出的各种有关语料库建设的理论、方法的基础上,紧密结合语料库开发的具体实例,全面、系统地研究了与语料库建设有关的理论与实践问题,讨论了如何为语言学研究的需要,建设语料库。 语料库是为某一个或多个应用而专门收集的、有一定结构的、有代表性的、可以被计算机程序检索的、具有一定规模的语料的集合。 语料库系统是以语料库为核心、包括计算机硬件、软件、语料库用户、语料采集和加工规则、语料库管理和应用程序的一个完整系统,其各部分互相影响、互相制约,共同决定语料库的质量、价值、应用水平。语料库系统这一概念的提出,有助于语料库建设时综合考虑有关的各方面的问题,形成一个有机的整体,从而提高语料库的质量和开发效率。 大型语料库的开发是一项软件工程,开发过程应遵循软件工程的一般原则和方法,但又要考虑自身的特点,故可以称为“语料库工程”。语料库工程的生命周期可以划分为7个阶段:语料库规划阶段、需求分析阶段、语料库设计阶段、语料采集阶段、语料库实现阶段、语料库标注阶段、语料库使用和维护阶段。 大型平衡语料库具有语料真实性、样本有限性、语料库代表性、库结构的平衡性等特点。语料真实性是语料库的立足之本,样本有限性是语料库不可回避的问题,代表性是语料库追求的目标,库结构的平衡性是达到目标的手段。 语料流是因特网上某一个或某几个站点源源不断产生的所有言语。当它流经监控程序时,监控程序获取可能需要的信息并保存起来,供后继的相关研究使用。可以根据需要,决定语料流中的语料是否需要长期保存。语料流的这一工作机制与人的大脑学习新知识、发现新知识的原理非常相似。基于语料流的监控语料库的建设,对于语言新现象的发现、报告有实际应用价值。 语料库的规范化是实现语料库的共享,减少语料库重复开发的关键;语料库的元数据规范化是语料库规范化工作中比较容易实现的一步,可以率先执行。语料库的元数据项可以分为六大类:语料知识版权信息、语料创建者背景信息、语料载体发行信息、语料内容信息、语料采集信息、语料库管理信息。 语料库标注的7条一般原则是:原始语料和标记符号的数据独立性原则、语料库的公开性原则、语料标注的通用性原则、语料标注的折衷性原则、语料标注的一致性原则、标注符号的确定性原则、用户知情权原则。 语料库标注过程中应该处理好以下几个关系:详细标注和简单标注的关系、通用性和专用性的关系、原则性和灵活性的关系、绝对性和模糊性的关系。 HNC理论建立了概念语义网络,可以用来描述词汇的语义,描述词汇之间的概念联想脉络。研究HNC概念表达式的形式化定义,旨在为语料库的自动语义标注建立语义知识表示体系,实现语义标注附码的形式化,实现语义的可计算性。 语料库应用工具软件的开发,能大力促进基于语料库的语言学研究,是语料库研究的一个重要内容,应该重视这方面的研究。

【Abstract】 The present paper is a study of corpus proper. It is based on linguistic theory and principles of software engineering and database. With the help of theory and methods of other related subjects and previous research findings, the paper analyzes some famous corpora, examines some academic and practical issues related to corpus construction and discusses how to construct corpus for the study of linguistics.Corpus is a representative collection of linguistic material with some kind of structure for application. It is large enough and machine-readable.The core of a corpus system is corpus. It also includes hardware, software, users of corpus, and the rules of collecting and processing linguistic material. Different parts of a corpus system affect and restrict one another. They work together to determine the quality and worthiness of corpus.The development of a big corpus can be regarded as a software engineering; therefore, it should follow the principles and methods of software engineering. However, it also has its own special features. So it can be called Corpus Engineering. The life cycle of a corpus engineering can be divided into seven phases: the planning phase, the needs analysis phase, designing phase, linguistic material collection phase, realizing phase, annotating phase, and using and maintenance phase.Balanced corpora have the following characteristics: authenticity of linguistic material, finity of the number of samples, representativeness of corpus, and balance of structure. The authenticity of linguistic material is the basis, the finity of sample is the reality, and the representativeness of corpus is the goal, while the balance is the means to realize the goal.The stream of linguistic material is all linguistic material produced continually from one or several web sites on the Internet. When it passes through monitor programs, the programs draw out useful information from it. Whether the stream should be stored depends on the need. The mechanism of linguistic material stream is similar to that of man’s brain. The construction of monitor corpus based on the stream of linguistic material is useful for the finding of new linguistic phenomenon.The normalization of corpus is the key to make corpora sharable, thus to reduce the repetition of corpora. The jiormalization of corpus meta-data is an easier step and can be done first. The corpus meta-data can be divided into six classes: information about copyright, information about background of linguistic material creator, information about medium of linguistic information, information aboutthe content of linguistic material, information about collecting linguistic material, and information about management of linguistic material.The general rules for annotating corpus are data independence of original linguistic material and annotating symbols, publicity of corpus, generality, compromise, consistency, correctness of annotation symbols, and user’s rights to know all about corpus.In the process of annotating corpus, the following relations are important: detailed and simple; general and specific; principled and flexible, absolute and indefinite.HNC theory sets up a network of concept. It can be used to describe the meaning of words and the associations between different’ concepts. The study of formalized definition of HNC concept expression aims at building a system of word semantic knowledge for automatic semantic annotation so that the formalization of semantic annotation symbols and calculability of the meaning of words can be realized.The development of application software for corpus can promote the corpus-based linguistic study. It is an important aspect of corpus study. So we should attach more importance to it.

节点文献中: