节点文献

应用条件随机场进行汉语词法分析、语块分析研究

The Research of Applying Conditional Random Fields to Chinese Lexical Analysis and Chunk Parsing

【作者】 罗恒

【导师】 王继曾;

【作者基本信息】 兰州理工大学 , 计算机应用技术, 2006, 硕士

【摘要】 介绍了词法分析、句法分析在自然语言理解研究中基础的、决定性的重要地位。针对目前词法分析、句法分析研究热点集中在基于规则方法和基于统计方法的联合应用,介绍了最大熵理论和最大熵理论对于自然语言理解研究的重要意义,并进一步介绍了条件随机场(以最大熵理论为驱动发展起来的一种用于对序列数据进行切分和标记的概率框架)。提出了应用条件随机场来构建统一的汉语词法分析。以往应用条件随机场进行汉语分词时,将分词转化为对汉字的标注。提出了使用词图作为基础的标记序列来完成汉语的词法分析,这样充分利用了现有的词典资源,在特征架的选择时也可以方便地融合语言知识。最后进一步讨论了将条件随机场应用到汉语语块分析之中。提出了未来关于应用条件随机场构建汉语词法语块分析模型的初步构想。

【Abstract】 This dissertation introduces the research of lexical analysis and syntax parsing is important, crucial and fundamental in the research of natural language understanding. According presently the tendency of methods that integrate statistics-based and rule-base methods, this paper introduces the rules of Maximum Entropy and the significance of it on natural language understanding research. Furthermore, this dissertation discusses the definition and parameter estimate of Condition Random Fields. CRFs are probabilistic models for segmenting and labeling sequence data and heavily motivated by the principle of maximum entropy. Then this dissertation presents a unified approach for Chinese lexical analysis using Conditional Random Fields. Precious applications applying conditional random fields to Chinese words segmentation convert segmentation to character-based Begin/Inside tagging. This dissertation presents using the words lattice as the fundamental sequence to be tagged to achieve Chinese lexical analysis. Then the lexicon can be used efficiently, and language knowledge can be integrated easily in feature template selecting. This dissertation also discusses applying Conditional Random Fields to Chinese Chunk Parsing and our future works.

  • 【分类号】TP391.1
  • 【被引频次】3
  • 【下载频次】397
节点文献中: 

本文链接的文献网络图示:

本文的引文网络