节点文献
基于隐马尔科夫模型的古汉语词性标注
Part-of-speech Tagging of Classical Chinese Based on Hidden Markovian Model
【摘要】 古汉语在语法和形态上与现代汉语有着本质的区别。从统计的角度出发,首先为古汉语设计一个标记集,将隐马尔可夫模型(HMM)与维特比算法相结合,以此对古汉语进行词性标注。通过对传统方法的改进,最终bigram模型和trigram模型的标注准确率分别提高到94.9%和96.5%,同时未登录词的标注精度也有显著提高。该方法应用于古汉语词性标注中,能根据古汉语的特点有效提高标注精度,并且在古汉语机器翻译等领域有广泛应用。
【Abstract】 Classical Chinese is essentially different from modern Chinese in grammar and form. From a statistical point of view, a tag set is designed for classical Chinese firstly, then Hidden Markovian Model(HMM) and Viterbi algorithm are used to tag part-of-speech in classical Chinese. The accuracies of bigram model and trigram model are improved to 94.9% and 96.5% respectively compared to traditional method, and the accuracy of unknown words is also improved significantly. This method can effectively improve the accuracy of part-of-speech tagging according to the characteristics of classical Chinese, and has wide applications in the field of machine translation of classical Chinese.
【Key words】 part-of-speech tagging; classical Chinese; Hidden Markovian model;
- 【文献出处】 微型电脑应用 ,Microcomputer Applications , 编辑部邮箱 ,2020年05期
- 【分类号】TP391.1
- 【被引频次】3
- 【下载频次】229