节点文献
SEBERT:基于语义信息的疫情文本命名实体识别
SEBERT: Semantic Information Based Named Entity Recognition for COVID-19 Text
【摘要】 近年来,新冠疫情持续爆发,如何快速处理疫情文本数据成为重要挑战。目前该领域公开的数据集相对较少,因此该文爬取了21个省级行政区的疫情公告并进行标注,构建了新冠疫情文本实体抽取数据集。由于数据集内实体较为密集,长实体过多,目前融入词汇信息的方法仍然不能有效地处理具有复杂结构的实体。该文根据疫情文本数据集的特点提出了一种基于语义信息的命名实体识别模型,通过将语义信息融入字向量增强模型对长实体的检测能力,并通过混合loss提高模型的学习能力。实验结果表明,该文提出的模型不仅适用于疫情文本数据,对比基线模型在4个中文数据集上F1值均有较高的提升。
【Abstract】 With the outbreak of COVID-19,how to quickly process epidemic text data has become an important challenge. This paper constructs a text entity extraction dataset of COVID-19 by crawling and labeling epidemic announcements from 21 provincial administrative regions. Because the entities in the data set are relatively dense and there are too many long entities, a named entity recognition model based on semantic information is further proposed according to the characteristics of the epidemic text data set.The semantic information is integrated into the word vector to enhance the detection ability of the model for long entities,and the learning ability of the model is improved by mixed loss. The experimental results showed that the model proposed in this paper significantly improves the F1 values of the four Chinese data sets compared with the baseline model.
【Key words】 named entity recognition; epidemic text; semantic information; mixing loss;
- 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2026年04期
- 【分类号】TP391.1
- 【下载频次】32