节点文献

结合学术网络与内容信息的文献语义表示方法研究

A Paper Semantic Representation Method Incorporating Academic Network with Content Information

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 石斌王昊李晓敏周抒

【Author】 Shi Bin;Wang Hao;Li Xiaomin;Zhou Shu;School of Information Management, Nanjing University;Jiangsu Key Laboratory of Data Engineering and Knowledge Service;

【通讯作者】 王昊;

【机构】 南京大学信息管理学院江苏省数据工程与知识服务重点实验室

【摘要】 随着科研工作者人数的不断增加,科技论文的发表数量呈现快速增长的趋势。面对海量的科技论文,文献的归档、录入和分析工作变得越发繁重。当前,针对文献的分类模型主要关注论文的内容信息,而忽略了论文相关的关联信息。为此,本文提出一种融合内容信息与学术网络的论文表征模型PAITKG (paper analysis by incorporating text and knowledge graph),引入知识图谱嵌入技术对文献的多重关联信息进行表征,采用Adapter微调的SciBERT提取内容特征,并将二者融合。在训练过程中,本文改进了动态对抗损失函数来引导模型更好地关注错误结果,并将该方法在数字人文和多模态学习两个领域的文献数据集上进行实验。在科技文献的学科多标签分类任务上,PAITKG比Baselines有显著改善,很好地提高了分类精度。除此以外,通过上游任务的学习,PAITKG的表征获得了更广泛的应用,在没有任何额外训练的情况下,本文模型提取的特征向量能够较好地应用于主题聚类、学者推荐等分析任务。研究结果表明,PAITKG通过构建并表征论文的学术网络,有效融合了文献的关联信息,提高了对文献数据的理解能力,而且其学习到的表征具有优秀的泛化潜力,能够应用于各种文献分析工作。

【Abstract】 With the increasing number of scientific research workers, the publication of scientific and technological papers published has increased rapidly, making the work of archiving, inputting, and analyzing documents increasingly burdensome. Most of the classification models focus on the content information of the paper, ignoring the relevant information.To solve this problem, this study proposes a paper representation model called PAITKG, which integrates content information and academic networks. Knowledge graph embedding technology is introduced to characterize multiple relationship patterns of literature; SciBERT, which is fine-tuned by Adapter, is used to extract content features and integrate the two. In the training process, this study improves the dynamic counter loss function to guide the model to pay more attention to error results. It applies this method to literature classification and analysis in the field of digital humanities. In the multilabel classification of scientific and technological literature, PAITKG showed significant improvement compared with the baselines, which greatly improved the classification accuracy. In addition, the representation of PAITKG has been more widely applied through the learning of upstream tasks. Without any additional training, the feature vectors extracted by the model can be applied to analysis tasks such as topic clustering and scholar recommendation. The experiments show that PAITKG can effectively integrate the associated literature information and improve the understanding of literature data by constructing and characterizing the academic networks of papers. Moreover, the representations learned by PAITKG have excellent generalization potential and can be applied to various literature analysis work.

【基金】 国家自然科学基金项目“关联数据驱动下我国非遗文本的语义解析与人文计算研究”(72074108);江苏省图书馆学会课题“江苏省公共文化服务适老化内容体系与目标路径研究”(22YB056)
  • 【文献出处】 情报学报 ,Journal of the China Society for Scientific and Technical Information , 编辑部邮箱 ,2025年02期
  • 【分类号】TP391.1
  • 【下载频次】157
节点文献中: 

本文链接的文献网络图示:

本文的引文网络