节点文献
构建大规模的汉语语块库
Build a large scale Chinese Functional Chunk Bank
【Author】 ZHOU QiangState Key Laboratory of IntelligentTechnology and Systems, Dept. ofComputer Science, Tsinghua.UniversityBeijing 100084ZHANG WeidongDepartment of ChineseLanguage & LiteraturePeking University, Beijing,100871REN HaiboInternational Cultural ExchangeCollegeShanghai Teachers University Shanghai 200234
【机构】 清华大学计算机系智能技术与系统国家重点实验室; 北京大学中文系; 上海师范大学国际文化交流学院;
【摘要】 本文介绍了构建200万字的汉语语块库的主要工作,包括设计语块标注体系、总结语块标注规范和协调语块加工流程等,分析了我们的标注体系与英语的CONLL-2000语块任务的主要差异,并提出了对现有标注体系的进一步理论思考和在现有语块库上的一些应用设想.
【Abstract】 In this paper, we firstly introduce some essential issues in the construction of a chunk bank with 2,000,000 Chinese characters, including functional chunk annotation schema, tagging specification and processing procedure. Then, we analyze the main difference of our annotation schema with CONLL-2000 shared chunking task, and propose some further theoretical thoughts of the current annotation schema and some application tentative based on the current chunk bank.
- 【会议录名称】 自然语言理解与机器翻译——全国第六届计算语言学联合学术会议论文集
- 【会议名称】全国第六届计算语言学联合学术会议
- 【会议时间】2001-08
- 【会议地点】中国山西
- 【分类号】H085
- 【主办单位】山西大学计算机系