节点文献
嵌入词库与多特征融合的冒犯性文本检测研究与实现
Research and Implementation of Offensive Text Detection with Embedded Lexicon and Multi-feature Fusion
【作者】 李娜;
【导师】 李邵梅;
【作者基本信息】 郑州大学 , 工程硕士(专业学位), 2024, 硕士
【摘要】 冒犯性文本是指包含辱骂、歧视、网络欺凌以及偏见言论等内容的文本。这类信息在网络上的恶意传播污染了网络环境,不利于构建和谐社会。因此,冒犯性文本检测技术应运而生。面向我国网络环境治理的需求,本文重点开展冒犯性中文文本检测技术研究。中文表达方式具有多样性,使得冒犯性中文文本的语义更加丰富。并且随着网络语言的快速演进,冒犯性文本中的新词频繁出现,给冒犯性文本检测带来更大的挑战,传统通用的文本分类方法在冒犯性中文文本检测中的准确率不高。为了提高冒犯性中文文本检测的准确率,本文首先利用新词发现算法构建了一个冒犯性词库;然后基于上述词库与多特征融合提出了冒犯性中文文本检测方法;最后基于上述提出的冒犯性中文文本检测方法,设计并实现了冒犯性文本检测系统。本文主要研究成果如下:(1)针对传统的新词发现算法在低频词识别上精确率不高的问题,提出一种结合规则统计和深度学习的冒犯性新词发现算法。首先根据冒犯性词汇特征,在使用K次互信息和邻接熵作为成词标准的基础上引入词性信息,并加入过滤规则进行筛选,获取初步候选词;然后利用初步候选词自动标注原始语料,并将基于依存句法树的位置向量以及CNN获取的局部特征引入Transformer-CRF模型中进一步识别新词;最后本文基于上述方法发现的冒犯性新词构建冒犯性词库,为后续的冒犯性中文文本检测提供支撑。实验结果表明,本文提出的方法相比于Transformer-CRF模型,新词发现的精确率提升了 1.9%。(2)针对部分冒犯性文本表达比较隐晦,字面特征不明显,检测难的问题,提出一种基于嵌入词库和多特征融合的冒犯性中文文本检测方法。该方法首先利用SoftLexicon将第一个研究点中自建冒犯性词库中的词汇引入WoBERT词向量中,增强WoBERT词向量的表征能力;然后将引入冒犯性词汇信息的WoBERT词向量和基于ALBERT提取的字向量进行拼接,再经过注意力机制层对字词融合向量做进一步的特征提取;最后将注意力机制的输出与基于ALBERT提取的句向量进行多维度特征融合,并基于上述字、词、句三级的融合特征进行冒犯性文本检测。实验结果表明,在基于COLDataset扩充的数据集上,本文提出的模型相比于基于BERT的检测模型,在准确率、召回率、F1值上分别提高了 1.72%、1.78%、1.71%。(3)基于本文的冒犯性中文文本检测方法,采用Flask框架设计并实现了一个冒犯性中文文本检测系统。该系统不仅支持用户输入单条文本进行冒犯性检测,同时也支持输入批量文本进行冒犯性检测。此外,系统支持从界面上查看检测结果。
【Abstract】 Offensive text refers to texts containing insults,discrimination,cyberbullying,and biased remarks,among other content.The malicious spread of such information pollutes the online environment,hindering the construction of a harmonious society.Therefore,offensive text detection technology has emerged.In response to the needs of governing the online environment in China,this thesis has focused on the research of offensive Chinese text detection technology.The diversity of expression in Chinese makes the semantics of offensive Chinese text richer.With the rapid evolution of internet language,new words frequently appear in offensive text,posing greater challenges to offensive text detection.Traditional general text classification methods have low accuracy in detecting offensive Chinese text.To enhance the accuracy of offensive Chinese text detection,an offensive lexicon has been constructed using a new word discovery algorithm.Then,based on the lexicon and multi-feature fusion,a method for detecting offensive Chinese text has been proposed.Finally,building upon the aforementioned method,an offensive text detection system has been designed and implemented.The main research achievements of this thesis are as follows:(1)In response to the low precision issue in identifying low-frequency words using traditional new word discovery algorithms,a novel offensive new word discovery algorithm has been proposed,which combines rule-based statistics and deep learning.Firstly,based on offensive vocabulary features,part-of-speech information is introduced on the basis of using K-means mutual information and critical entropy as word formation standards,and filtering rules are added for screening to obtain initial candidate words.Then,the initial candidate words are automatically annotated in the original corpus,and the position vectors based on dependency syntax trees and local features obtained by CNN are introduced into the Transformer-CRF model for further identification of new words.Finally,based on the aforementioned methods,this thesis constructed an offensive lexicon using the offensive new words discovered,providing support for subsequent offensive Chinese text detection.Experimental results show that compared to the Transformer-CRF model,the proposed method improves the precision of new word discovery by 1.9%.(2)In response to the challenge of detecting offensive texts with obscure expressions and less obvious literal features,a method for detecting offensive Chinese texts based on embedding word libraries and integrating multiple features has been proposed.This method first utilizes SoftLexicon to introduce the vocabulary from the self-built offensive lexicon in the first research point into the WoBERT word vectors,enhancing the representation ability of WoBERT word vectors.Then,the WoBERT word vectors with introduced offensive vocabulary information are concatenated with the character vectors extracted based on ALBERT,and further feature extraction is performed on the fused character-word vectors through the attention mechanism layer.Finally,the output of the attention mechanism is multidimensionally fused with the sentence vectors extracted based on ALBERT,and offensive text detection is performed based on the fused features at the character,word,and sentence levels.Experimental results indicate that on the dataset expanded based on COLDataset,the proposed model in this thesis achieved improvements of 1.72%,1.78%,and 1.71%in accuracy,recall,and F1 score respectively compared to the detection model based on BERT.(3)Based on the offensive Chinese text detection method proposed in this thesis,a detection system for offensive Chinese text has been designed and implemented using the Flask framework.This system not only supports users in entering individual texts for offensive detection,but also allows for the input of batches of texts for offensive detection.Additionally,the system enables users to view the detection results through the interface.
【Key words】 New Word Discovery; Dependency Syntax Tree; Offensive Lexicon Construction; Offensive Text Detection; Multi-Feature Fusion;
- 【网络出版投稿人】 郑州大学 【网络出版年期】2026年 06期
- 【分类号】TP391.1