节点文献

基于语义依存关系的汉语语料库的构建

On Construction of a Chinese Corpus Based on Semantic Dependency Relations

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 尤昉李涓子王作英

【Author】 YOU Fang 1,LI Juan-zi 2,WANG Zuo-ying 1 (1.Dept.of Electronics Engineering,Tsinghua University,Beijing 100084,China 2.Dept.of Computer Science Technology,Tsinghua University,Beijing 100084,China)

【机构】 清华大学电子工程系清华大学计算机科学与技术系清华大学电子工程系 北京100084北京100084北京100084

【摘要】 语料库是自然语言处理中用于知识获取的重要资源。本文以句子理解为出发点 ,讨论了在设计和建设一个基于语义依存关系的汉语大规模语料库过程中的几个基础问题 ,包括 :标注体系的选择、标注关系集的确定 ,标注工具的设计 ,以及标注过程中的质量控制。该语料库设计规模 10 0万词次 ,利用 70个语义、句法依存关系 ,在已具有语义类标记的语料上进一步标注句子的语义结构。其突出特点在于将《知网》语义关系体系的研究成果和具体语言应用相结合 ,对实际语言环境中词与词之间的依存关系进行了有效的描述 ,它的建成将为句子理解或基于内容的信息检索等应用提供更强大的知识库支持。

【Abstract】 Corpora are important resources for knowledge acquisition in the field of natural language processing.For the purpose of sentence understanding,we are constructing a Chinese large-scale-corpus based on semantic dependency relations.This paper introduces the tagging formalisms we adopt,the tagging set we choose,the tagging tool we develop,and the method we use to guarantee the good consistency of tagging.The corpus under discussion is at a scale of 1 million words.Each sentence in the corpus,which already had annotations of sense,is further tagged with its semantic structure using 70 semantic-dependency-relations.The highlight of this corpus is its ability to effectively describe various relations between Chinese words.All of these profited from using <HowNet> for reference and the combination with specific use of language.The construction of this corpus can definitely provide more knowledge supports for sentence understanding,content-based information retrieval,and so on.

  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2003年01期
  • 【分类号】TP391.1
  • 【被引频次】50
  • 【下载频次】1075
节点文献中: