节点文献

基于深度学习的分子表示学习模型开发及在药物发现中的应用

Development of Deep Learning-Based Molecular Representation Learning Models and Their Applications in Drug Discovery

【作者】 安康;

【导师】 申兆岩;

【作者基本信息】 山东大学 , 计算机技术(专业学位), 2025, 硕士

【摘要】 随着癌症等重大疾病的治疗需求日益增加,药物研发的挑战越来越复杂,特别是小分子抗癌药物的发现。传统的药物筛选方法,如高通量筛选和计算机辅助分子对接,往往面临高成本和低效率的限制。近年来,基于深度学习的分子表示学习方法因其在处理复杂数据中的强大能力,已成为药物发现领域的重要研究方向。然而,现有的药物性质预测方法在准确性、泛化能力和可解释性方面仍然存在挑战,限制了其在实际应用中的广泛推广。本研究旨在构建基于深度学习的分子表示学习模型,提升分子药物性质预测的准确性,并应用于小分子抗癌药物的发现任务。本论文主要包括两个核心研究内容:(1)基于分子图和分子指纹的分子表示学习模型用于SOS1抑制剂筛选与活性预测:Son of Sevenlesshomolog 1(SOS1)作为RAS信号通路中的关键鸟嘌呤核苷酸交换因子,广泛参与肺癌、胰腺癌和结直肠癌等恶性肿瘤的异常激活,因此成为抗癌药物研发的重要靶点。传统的SOS1抑制剂筛选方法面临成本高且效率低的问题。本研究提出了结合多头注意力机制增强的图神经网络、摩根指纹和药效团ErG指纹的分子表示学习模型,通过多尺度信息融合策略构建了更加全面的分子表征方式,提升了 SOS1抑制剂活性预测的精度和泛化能力。实验结果表明,该模型在多个药物性质预测基准数据集上表现优异,并且在真实世界的药物筛选任务中展现了较高的应用潜力。(2)基于深度学习的分子图嵌入器和分子指纹嵌入器的开发与验证:本研究采用多尺度信息融合策略,结合分子指纹和分子图结构,创新的提出了分子指纹嵌入器FP-Encoder和分子图嵌入器TransGAT,优化了分子表示学习模型的性能,提高了模型在药物性质预测任务中的精度。该模型在9个药物性质预测基准数据集上均取得了最优预测性能,验证了所提出方法的有效性和泛化能力。此外,研究通过消融实验系统地评估了不同嵌入器对模型性能的贡献,深入探讨了 TransGAT中的全局特征提取模块和局部特征提取模块的作用,并对FP-Encoder与传统MLP方法在分子指纹处理中的性能差异进行了深入分析。模型还引入了注意力机制进行可解释性分析,解析了模型在预测过程中关注的关键分子子结构,从而提高了模型的透明度和可信度。本研究提出的两种深度学习驱动的分子表示学习方法,能够提高SOS 1抑制剂筛选、癌症靶点生物活性预测和药物性质预测任务中的预测准确性,为基于AI的药物发现提供了新的计算工具。通过多尺度信息融合,模型能够有效整合分子指纹与分子图的特征,提升对不同分子结构的适应性,进一步推动了深度学习在药物研发、分子筛选和优化中的实际应用。

【Abstract】 With the increasing demand for the treatment of major diseases such as cancer,the challenges in drug development have become progressively more complex,particularly in the discovery of small-molecule anticancer drugs.Traditional drug screening methods,such as high-throughput screening and computer-aided molecular docking,are often constrained by high costs and low efficiency.In recent years,deep learning-based molecular representation learning methods have become a crucial research direction in drug discovery due to their powerful capability in handling complex data.However,existing drug property prediction methods still face challenges in terms of accuracy,generalization,and interpretability,limiting their widespread application in real-world scenarios.This study aims to develop a deep learning-based molecular representation learning model to enhance the accuracy of molecular drug property prediction and apply it to small-molecule anticancer drug discovery tasks.The main contributions of this dissertation are as follows:(1)Deep Learning-Based Molecular Representation Learning Model for SOS1 Inhibitor Screening and Activity Prediction:Son of Sevenless homolog 1(SOS1)is a key guanine nucleotide exchange factor in the RAS signaling pathway,which is abnormally activated in various malignancies such as lung cancer,pancreatic cancer,and colorectal cancer.As a result,SOS1 has become an important target in anticancer drug development.Traditional SOS1 inhibitor screening methods face challenges of high costs and low efficiency.This study proposes a molecular representation learning model that integrates a multi-head attentionenhanced graph neural network,Morgan fingerprints,and ErG pharmacophore fingerprints.Through a multi-scale information fusion strategy,a more comprehensive molecular representation is constructed,which improves the accuracy and generalization of SOS1 inhibitor activity prediction.Experimental results show that the model performs excellently on several drug property prediction benchmark datasets and demonstrates high application potential in real-world drug screening tasks.(2)Development and Validation of Deep Learning-Based Molecular Graph and Fingerprint Embedders:This study adopts a multi-scale information fusion strategy that combines molecular fingerprints and molecular graph structures.It introduces innovative molecular fingerprint embedder(FP-Encoder)and molecular graph embedder(TransGAT),which optimize the performance of molecular representation learning models and enhance their accuracy in drug property prediction tasks.The model achieves optimal prediction performance across nine drug property prediction benchmark datasets,validating the effectiveness and generalization ability of the proposed method.Furthermore,ablation experiments systematically evaluate the contribution of different embedders to model performance,exploring the roles of the global feature extraction module and the local feature extraction module in TransGAT.The study also conducts an in-depth analysis of the performance differences between FP-Encoder and traditional Multilayer Perceptron methods in processing molecular fingerprints.Additionally,the model incorporates attention mechanisms for interpretability analysis,highlighting key molecular substructures that the model focuses on during prediction,thereby improving the model’s transparency and credibility.The two deep learning-driven molecular representation learning methods proposed in this study enhance the prediction accuracy of SOS1 inhibitor screening,cancer target bioactivity prediction,and drug property prediction tasks,providing new computational tools for AI-based drug discovery.Through multi-scale information fusion,the model effectively integrates molecular fingerprint and molecular graph features,improving its adaptability to various molecular structures and further advancing the practical application of deep learning in drug development,molecular screening,and optimization.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2026年 06期
  • 【分类号】R91;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络