节点文献
基于图卷积神经网络的古汉语分词研究
Ancient Chinese Word Segmentation Based on Graph Convolutional Neural Network
【摘要】 古汉语的语法有省略、语序倒置的特点,词法有词类活用、代词名词丰富的特点,这些特点增加了古汉语分词的难度,并带来严重的out-of-vocabulary (OOV)问题。目前,深度学习方法已被广泛地应用在古汉语分词任务中并取得了成功,但是这些研究更关注的是如何提高分词效果,忽视了分词任务中的一大挑战,即OOV问题。因此,本文提出了一种基于图卷积神经网络的古汉语分词框架,通过结合预训练语言模型和图卷积神经网络,将外部知识融合到神经网络模型中来提高分词性能并缓解OOV问题。在《左传》《战国策》和《儒林外史》 3个古汉语分词数据集上的研究结果显示,本文模型提高了3个数据集的分词表现。进一步的研究分析证明,本文模型能够有效地融合词典和N-gram信息;特别是N-gram有助于缓解OOV问题。
【Abstract】 The syntax of ancient Chinese is characterized by the omission and inversion of word order, and morphology is characterized by the word-class shift and the abundance of pronouns and nouns. These features increase the difficulty of ancient Chinese word segmentation(CWS) and lead to the serious out-of-vocabulary(OOV) problem. Recently, deep learning methods have been widely used on ancient CWS tasks and achieved significant success. However, these works paid more attention to improving the performance of CWS and ignored the OOV issue, a major challenge in CWS. Therefore,we propose an ancient CWS framework that combines the pre-trained language model and the graph convolutional neural network, integrating external knowledge into the neural network model to relieve the OOV problem. The experimental results on three ancient Chinese CWS datasets(Zuo Zhuan, Stratagems of the Warring States, and The Scholars) demonstrate that our model improves the word segmentation performance of the three datasets. Further analysis illustrates that our model can effectively integrate lexicon and N-gram information. In particular, N-gram helps to alleviate the OOV problem.
【Key words】 ancient Chinese; Chinese word segmentation; graph convolutional neural network; pre-trained language model; BERT(bidirectional encoder representations from transformers);
- 【文献出处】 情报学报 ,Journal of the China Society for Scientific and Technical Information , 编辑部邮箱 ,2023年06期
- 【分类号】TP391.1;TP183
- 【下载频次】73