节点文献

汉语组块的定义和获取

Research on Definition and Acquisition of Chunk

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李素建刘群

【Author】 Li Sujian, Liu Qun Institute of Computational Linguistics, Peking University, Peking, China, 100871

【机构】 北京大学计算语言学研究所

【摘要】 组块是介于词语和句子之间的一种语言结构,目前还没有明确的定义。本文总结了当前对组块的各种研究,对汉语组块进行了定义。同时组块的获取和收集也是一项迫切的任务,由于不易直接获取到具有组块标注的语料,我们从现有树库中抽取组块。本文根据汉语特点提出了12种汉语组块类型,并根据这些组块类型和宾州大学中文树库短语类型的对应关系进行转化获得组块库。

【Abstract】 Chunk is a kind of linguistic structure between word and sentence, which isn’t defined definitely now. This paper summarizes various current researches on chunks, and defines what is a Chinese chunk. At the same time, the acquisition and collection of chunks are a hard but urgent work. Due to the difficulty of acquiring chunked corpus, we adopt the method of converting from Treebank available. According to the characteristics of Chinese. 12 Chinese chunk categories are proposed. Then our chunked corpus is obtained by extracting from Upenn Chinese Treebank.

【关键词】 组块组块语料库树库语法分析
【Key words】 ChunkChunked corpusTreebankSyntactic parsing
  • 【会议录名称】 语言计算与基于内容的文本处理——全国第七届计算语言学联合学术会议论文集
  • 【会议名称】全国第七届计算语言学联合学术会议
  • 【会议时间】2003-08
  • 【会议地点】中国哈尔滨
  • 【分类号】TP391.1
  • 【主办单位】哈尔滨工业大学计算机科学与技术学院、清华大学智能技术与系统国家重点实验室
节点文献中: 

本文链接的文献网络图示:

本文的引文网络