节点文献
结合双预训练语言模型的中文文本分类模型
Chinese text classification model based on dual pre-trained language model
【摘要】 针对Word2Vec等模型所表示的词向量存在语义模糊从而导致的特征稀疏问题,提出一种结合自编码和广义自回归预训练语言模型的文本分类方法。首先,分别通过BERT、XLNet对文本进行特征表示,提取一词多义、词语位置及词间联系等语义特征;再分别通过双向长短期记忆网络(BiLSTM)充分提取上下文特征,最后分别使用自注意力机制(Self_Attention)和层归一化(Layer Normalization)实现语义增强,并将两通道文本向量进行特征融合,获取更接近原文的语义特征,提升文本分类效果。将提出的文本分类模型与多个深度学习模型在3个数据集上进行对比,实验结果表明,相较于基于传统的Word2Vec以及BERT、XLNet词向量表示的文本分类模型,改进模型获得更高的准确率和F1值,证明了改进模型的分类有效性。
【Abstract】 To solve the problem of sparse features caused by Word2Vec models, a text classification method based on autocoding and generalized autoregressive pretrained language model is proposed. Firstly, BERT and XLNet are used to represent features such as polysemy, word location and relationship of the text respectively. Then, context features are extracted through Bi-directional Long Short-Term Memory(BiLSTM). Finally, features through self-attention and layer normalization of the two channel are fused to obtain the features closer to the original text and improve the effect of text classification. The proposed text classification model is compared with several deep learning models on three datasets. Experimental results show that compared with traditional text classification models based on Word2Vec, BERT and XLNet word vector representation, the improved model achieves higher accuracy and F1, which proves the validity of the improved model.
【Key words】 pre-trained language model; Bi-directional Long Short-Term Memory; self-attention mechanism; layer normalization;
- 【文献出处】 智能计算机与应用 ,Intelligent Computer and Applications , 编辑部邮箱 ,2023年07期
- 【分类号】TP391.1
- 【下载频次】29