节点文献
一种基于NA假设的训练数据自动构造方法
Training Data Acquisition Based on NA Assumption
【摘要】 为减轻人工标注训练语料库面临的瓶颈问题,提出了一种基于 N A 假设带标训练语料库的自动构造方法·为了检验该方法的有效性,将自动获取的带标训练语料库用于词性标注应用中,2 万词次的开放性测试结果的准确率为93 .1 % ,其中词性兼类消歧准确率为79 .3 % ,未登录词词性确定准确率为88 % ·
【Abstract】 An approach to training data acquisition based on NA assumption from a raw corpus was presented. Using the training data trains parameters of the grammatical tagging model. The precision of grammatical tagging on the open test set is 93%. The precision of multi POS′s word disambiguation is 79 3%. The precision of determining grammatical category of unknown word is 88%.
【关键词】 N元语法;
词性标注;
自然语言理解;
【Key words】 n gram model; grammatical tagging; natural language understanding.;
【Key words】 n gram model; grammatical tagging; natural language understanding.;
【基金】 国家自然科学基金
- 【文献出处】 东北大学学报 ,JOURNAL OF NORTHEASTERN UNIVERSITY , 编辑部邮箱 ,1999年04期
- 【分类号】TP391.2
- 【被引频次】5
- 【下载频次】63