节点文献

汉语分词及词性标注自动校验方法研究

Research of Verifying Method of Chinese Word Segmentation and Part-of-speech Tagging

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 钱揖丽张虎

【Author】 QIANYili ZHANG Hu (Shanxi University, Taiyuan 030006, China)

【机构】 山西大学计算机科学系

【摘要】 大规模的标注语料库是语料库语言学发展的重要基础。随着许多科学研究的进一步开展,我们对语料的加工质量提出了更高的要求。本文采用基于上下文搭配的规则和统计相结合的自动校验方法,对机器切分标注语料进行处理,并把自动校验过程中获取的信息,应用于语料库的构建,即采用滚动式的方法,建立大规模的、具有更高加工质量的标注语料库。

【Abstract】 The large-scale tagged corpus is the important basis of the development of corpus linguistics. Many reaseaches request corpora with higher processing quality. This paper presents a verifying method based on rules and statistic. This method obtains informations from the corpora’s verifying course, and then applies these informations to the building of corpus. We use rolling method, build large-scale Chinese corpus, and we have obtained higher processing quality.

  • 【会议录名称】 第一届学生计算语言学研讨会论文集
  • 【会议名称】第一届学生计算语言学研讨会
  • 【会议时间】2002-08
  • 【分类号】H085
  • 【主办单位】中国中文信息学会
节点文献中: