节点文献
结合学术网络与内容信息的文献语义表示方法研究
A Paper Semantic Representation Method Incorporating Academic Network with Content Information
【摘要】 随着科研工作者人数的不断增加,科技论文的发表数量呈现快速增长的趋势。面对海量的科技论文,文献的归档、录入和分析工作变得越发繁重。当前,针对文献的分类模型主要关注论文的内容信息,而忽略了论文相关的关联信息。为此,本文提出一种融合内容信息与学术网络的论文表征模型PAITKG (paper analysis by incorporating text and knowledge graph),引入知识图谱嵌入技术对文献的多重关联信息进行表征,采用Adapter微调的SciBERT提取内容特征,并将二者融合。在训练过程中,本文改进了动态对抗损失函数来引导模型更好地关注错误结果,并将该方法在数字人文和多模态学习两个领域的文献数据集上进行实验。在科技文献的学科多标签分类任务上,PAITKG比Baselines有显著改善,很好地提高了分类精度。除此以外,通过上游任务的学习,PAITKG的表征获得了更广泛的应用,在没有任何额外训练的情况下,本文模型提取的特征向量能够较好地应用于主题聚类、学者推荐等分析任务。研究结果表明,PAITKG通过构建并表征论文的学术网络,有效融合了文献的关联信息,提高了对文献数据的理解能力,而且其学习到的表征具有优秀的泛化潜力,能够应用于各种文献分析工作。
【Abstract】 With the increasing number of scientific research workers, the publication of scientific and technological papers published has increased rapidly, making the work of archiving, inputting, and analyzing documents increasingly burdensome. Most of the classification models focus on the content information of the paper, ignoring the relevant information.To solve this problem, this study proposes a paper representation model called PAITKG, which integrates content information and academic networks. Knowledge graph embedding technology is introduced to characterize multiple relationship patterns of literature; SciBERT, which is fine-tuned by Adapter, is used to extract content features and integrate the two. In the training process, this study improves the dynamic counter loss function to guide the model to pay more attention to error results. It applies this method to literature classification and analysis in the field of digital humanities. In the multilabel classification of scientific and technological literature, PAITKG showed significant improvement compared with the baselines, which greatly improved the classification accuracy. In addition, the representation of PAITKG has been more widely applied through the learning of upstream tasks. Without any additional training, the feature vectors extracted by the model can be applied to analysis tasks such as topic clustering and scholar recommendation. The experiments show that PAITKG can effectively integrate the associated literature information and improve the understanding of literature data by constructing and characterizing the academic networks of papers. Moreover, the representations learned by PAITKG have excellent generalization potential and can be applied to various literature analysis work.
【Key words】 paper represent; semantic representation; associated feature; knowledge graph embedding; RelaGraph;
- 【文献出处】 情报学报 ,Journal of the China Society for Scientific and Technical Information , 编辑部邮箱 ,2025年02期
- 【分类号】TP391.1
- 【下载频次】157