节点文献

基于语义分析模型的涉毒人员聊天文本挖掘

【作者】 张立;

【导师】 范馨月;

【作者基本信息】 贵州大学 , 数学, 2021, 硕士

【摘要】 对涉毒人员聊天文本进行语义分析,可从海量复杂的网络中快速精准地挖掘出涉毒人员并及时追踪调查。针对禁毒研判平台所采集到的实时数据进行有效数据选择,利用带有方言特色和特定语境下吸毒信息采集平台的吸毒人员聊天文本数据,以一般文本分类模型为出发点训练涉毒人员聊天文本,和基于上下文语义学习的BERT模型进行理论分析、实验验证,并根据存在的问题进行分析并作出相应的改进。可学习上下文语境的BERT模型,对聊天文本涉毒数据挖掘效果显著,在准确率、召回率和F1值均优于一般分类模型。具体的研究工作与成果如下:(1)通过学习分散式和分布式两类文本表示方法,从传统词向量模型开始,采用TF-IDF和贝叶斯分类模型来分析涉毒人员聊天文本数据,观察发现此类吸毒人员聊天文本数据,在不同语境下一词多义的字出现次数较多时,模型判别能力较差,进行文本正确分类工作存在困难,需要进行多义词消歧义。(2)为考虑上下文关系的影响,提出使用BERT模型。完成BERT模型预训练的微调,利用得到的最佳学习率进行文本分类工作。在测试文本中,BERT模型在准确率上高出贝叶斯模型7个百分点,缉毒文本分类任务总体优于一般文本分类模型。(3)分析BERT模型错判数据的内容结构,发现针对语句出现分散性敏感词时,模型判别能力不强,考虑在文字编码中添加敏感词的影响。借助敏感词库提取、输出文本敏感词,融入BERT预训练模型中,建立BERT-sen预训练模型,重新学习、输出具体场景中字的向量表示。针对性学习BERT模型错判语句后,在测试文本中,BERTsen预训练模型在准确率上高出BERT模型3个百分点,在学习多词义文本时比BERT模型更加敏感有效。

【Abstract】 Semantic analysis of the chat text of drug-related personnel can dig out drug-related personnel from the massive and complex network and then investigate them in time quickly and accurately.This paper makes effective data selection based on the real-time data collected by the anti-drug research and judgment platform,and uses the chat text data of drug-related personnel with dialect characteristics and the chat text in a specific context.The general text classification model is used as the starting point to train the chat text of drug-related personnel,and the Bert model based on contextual semantic learning is used for theoretical analysis and experimental verification,then the existing problems are analyzed and corresponding improvements are made.The BERT model,which can learn the context,has a significant effect on drug-related data mining of the chat text,and is better than the general classification model in accuracy,recall and F1 value.The specific research work and results of this paper are as follows:(1)By learning the decentralized and distributed text representation,using the traditional word vector model,TF-IDF and Bayesian classification model to analyze the chat text data of drug-related personnel,it is observed that in the chat text data of this kind of drug-related personnel,when there are many occurrences of polysemous words in different contexts,the model’s discrimination ability is poor,and it is difficult to classify the text correctly,and it is necessary to disambiguate polysemous words.(2)In order to consider the influence of context,the BERT model is proposed.Finishing the fine-tuning of the pre-training of the BERT model,and using the best learning rate obtained for the text classification.In the test text,the accuracy of the BERT model is 7 percentage points higher than that of the Bayesian model,and the anti-drug text classification task is generally better than the Bayesian model.(3)By analyzing the content structure of the misjudgment data of the BERT model,it is found that the model’s discriminant ability is not strong when there are scattered sensitive words in the sentence,so the influence of adding sensitive words in the text encoding is considered.With the help of the sensitive word database,text sensitive words are extracted and outputted,and then integrate them into the Bert pre-trained model.The BERT-sen pre-trained model is established to relearn and output the vector representation of words in specific scenes.After learning the wrong sentence of the BERT model,the accuracy of the BERT-sen pre-trained model is 3% higher than that of the BERT model in the test text.It is more sensitive and effective than the BERT model when learning multi-word text.

  • 【网络出版投稿人】 贵州大学
  • 【网络出版年期】2022年 05期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络