节点文献
汉语基本块规则的自动学习和扩展进化
Automatic learning and refinement algorithm for Chinese base chunk rules
【摘要】 为了从大规模标注语料库和词汇知识库支持下自动获取分层次、多粒度的规则描述知识,从汉语多词语基本块入手,提出一套完整处理方案。该方案从标注语料库中自动获取所有基于词类的基本块规则,通过设置规则置信度自动排除大量低可靠和无效规则。针对其中的高频低可靠规则,不断引入更多的内部词汇约束和外部语境限制知识,使之逐步进化为描述能力更强的结构化规则。同时提出一种预期精度指标对自动习得规则的描述能力进行了客观评价。实验结果表明:现有算法以16%的有效扩展规则覆盖了93%的标注正例,并使预期精度从51%提高到81%,显示了这套规则学习和评价方法的有效性。
【Abstract】 A method is presented to automatically learn and refine Chinese base chunk rules,using a large annotated corpus and a lexical knowledge base.After extracting all possible parts-of-speech-based rules from the annotated corpus,the system first prunes most of useless rules,and expands some low reliability rules with hierarchical knowledge from the internal lexical relationships and external contextual restrictions.The system then refines the rules into structural rules with stronger descriptive capabilities.A confidence score computation is used to evaluate rule reliability during the learning procedure,with an expected accuracy index to evaluate the descriptive capabilities of the refined rule base.Test results indicate that the algorithm can acquire about 16% of the useful expanded rules to cover 93% of the annotated positive examples and can improve the expected accuracy from 51% to 81%.
【Key words】 information processing; rule knowledge acquisition; base chunk; confident score analysis; restriction-based refinement; rule base evaluation;
- 【文献出处】 清华大学学报(自然科学版) ,Journal of Tsinghua University(Science and Technology) , 编辑部邮箱 ,2008年01期
- 【分类号】TP391.1
- 【被引频次】12
- 【下载频次】121