节点文献

基于多核融合和特征提取的药物—疾病关联预测方法研究

【作者】 王颖;

【导师】 刘金星;

【作者基本信息】 曲阜师范大学 , 计算机应用技术, 2024, 硕士

【摘要】 癌症已经成为中国高发疾病之一,找到治疗癌症所需的药物是至关重要的。然而,传统的药物开发需要花费大量的时间和金钱,识别药物与疾病之间的关联可以为生物学家进行下一步的湿实验提供有效的候选药物,因此,设计低成本且高效的计算模型识别药物与疾病之间的关联成为研究热点之一。尽管现有的模型已经取得不错的预测结果,但大多数都存在相似性信息融合方式和特征提取方式单一的问题。本文是在降低先验信息噪声的前提下,从多种相似性信息融合和多种提取方式出发,提出了四个基于多核融合和特征提取的药物-疾病关联预测方法。预测方法的提出为药物的开发提供了新的方向,具体的研究内容如下:(1)针对药物-疾病关联预测中特征提取方式过于单一的问题,提出了基于图表示学习和轻量级梯度提升机的方法(GRLGB)预测药物与疾病之间的关联。首先,将蛋白质信息引入异构网络。然后,使用不同的图表示学习方法分别从网络拓扑和生物知识角度提取药物和疾病的节点特征,使用了两种不同的方式处理生物知识。最后,使用轻量级梯度提升机预测潜在的药物-疾病关联。(2)针对相似性信息过于单一以及相似性融合方式过于简单的问题,提出了一种基于多种相似性和图卷积自动编码器的方法(MSGCA)预测药物-疾病关联。首先,使用基于中心核对齐的多核学习(CKA-MKL)算法融合多种药物相似性和多种疾病相似性,然后,通过使用线性邻域相似性改善经CKA-MKL融合的药物相似性和疾病相似性,同时使用加权K最近邻谱算法减少原始关联矩阵的稀疏性。最后,将处理后的矩阵输入异构网络,然后使用带注意力机制的图卷积自动编码器预测潜在的药物-疾病关联。(3)针对先验信息存在噪声和特征提取方式单一的问题,提出了一种基于堆叠去噪自动编码器和矩阵分解的方法(SDAEMF)预测药物-疾病关联。首先,使用重启的随机冲浪模型和计算移位正点互信息矩阵对药物和疾病相似性网络进行预处理,从而减少药物和疾病相似性网络的噪声。其次,使用矩阵分解挖掘原始关联矩阵的潜在药物和疾病特征,使用堆叠去噪自动编码器提取高阶低维特征表示。然后,将矩阵分解和堆叠去躁自动编码器得到的药物和疾病的特征矩阵分别串联。接下来,将串联后的药物特征矩阵和疾病特征矩阵分别使用核邻域相似性进行融合,从而挖掘不同相似性网络之间的非线性信息及相似性网络之间的邻域和非邻域的非线性信息。最后,使用图卷积自动编码器预测潜在的药物-疾病关联。(4)针对特征提取和融合方式过于单一以及过分依赖原始关联矩阵的问题,提出了基于堆叠去噪自动编码器和改进的扩散成分分析(clusDCA)方法(SDAEDCA)预测药物和疾病之间的关联。首先,使用基于熵的加权融合方法对药物和疾病相似性网络进行处理,从而根据贡献分配不同的权重。其次,使用堆叠去噪编码器和clusDCA提取药物和疾病相似性网络的高阶低维特征。然后,将矩阵分解和堆叠自动编码器得到的药物和疾病的特征矩阵分别串联。接下来,将药物和疾病特征矩阵分别使用相似性网络融合方法融合。最后,使用拉普拉斯正则化最小二乘法预测潜在的药物-疾病关联。本文提出的四种方法都是基于特征提取进行的预测,且均用于药物与疾病之间的关联预测,同时具有优越的预测性能。最终的实验结果表明,更先进的特征提取方法可以显著提升方法的预测效果,从而为药物的开发提供新的思路。

【Abstract】 Cancer has become one of the most common diseases in China,and finding the drugs needed to treat these cancers is crucial.However,traditional drug development requires a significant amount of time and money,and identifying the associations between drugs and diseases can provide effective candidate drugs for biologists to conduct the next wet experiment.Therefore,designing low-cost and efficient computational models to identify the associations between drugs and diseases has become one of the research hotspots.Although existing models have achieved good prediction results,most of them have problems with single feature fusion and feature extraction methods.Under the premise of reducing the prior information noise,the paper proposes four drug-disease associations prediction methods based on multi-kernel fusion and feature extraction from multiple similarity information fusion and multiple extraction methods.The proposal of prediction methods provides new solutions for drug discovery.The specific research content is as follows:(1)Aiming at the problem that the feature extraction method in drug-disease associations prediction is too single,a method based on graph representation learning and light gradient boosting machine(GRLGB)was proposed to predict the associations between drugs and diseases.First,protein was introduced into the heterogeneous network.Then,different graph learning methods were used to extract node features of drugs and diseases from the perspectives of network topology and biological knowledge.The method uses two different approaches to process biological knowledge.Finally,the light gradient boosting machine was used to predict potential drug-disease associations.(2)Aiming at the problems of single similarity information and simple similarity fusion methods,a method based on multiple similarities and graph convolutional autoencoder(MSGCA)was proposed to predict drug-disease associations.First,the centered kernel alignment-based multiple kernel learning algorithm(CKA-MKL)was used to fuse multiple drug and disease similarities,and then linear neighborhood similarity was used to improve the fused drug and disease similarities through CKA-MKL,while weighted K nearest neighbor profile was used to reduce the sparsity of associations matrix.Finally,the processed matrices were put into a heterogeneous network,and a graph convolutional autoencoder with an attention mechanism was used to predict potential drug-disease associations.(3)Aiming at the problems of prior information noise and single feature extraction method,a method based on stacked denoising autoencoders and matrix factorization(SDAEMF)was proposed to predict the associations between drugs and diseases.First,the drug and disease similarity networks were preprocessed using a restarted random surfing model and computing a shifted positive point mutual information matrix to reduce the noise of drug and disease similarity.Secondly,the latent drug and disease features of the original associations matrix were mined through matrix factorization,and stacked denoising autoencoder was used to extract high-order low-dimensional feature representations.Then,the extracted drug and disease feature matrices obtained from matrix factorization and stacked denoising autoencoder were concatenated,respectively.Next,the concatenated drug feature matrix and disease feature matrix were fused using kernel neighborhood similarity to mine the nonlinear information between different similarity networks and the nonlinear information of neighborhood and non-neighborhood of similarity networks.Finally,a graph convolutional autoencoder was used to predict drug-disease associations.(4)Aiming at the problem that feature extraction and feature fusion methods are too single and too dependent on the original associations matrix,a method(SDAEDCA)based on stacked denoising autoencoder and improved diffusion component analysis(clusDCA)was proposed to predict the associations between drugs and diseases.First,an entropy-based weighted fusion was used to preprocess the drug and disease similarity networks to assign different weights based on contributions.Secondly,a stacked denoising autoencoder and clusDCA were used to extract highorder low dimensional features of the drug and disease similarity networks,and then the obtained drug feature matrix and disease feature matrix from stacked denoising autoencoder and clusDCA are concatenated,respectively.Next,the drug and disease feature matrices were fused using similarity network fusion method,respectively.Finally,the Laplacian Regularized Least Squares was used to predict latent drug-disease associations.The four methods proposed in the article are based on feature extraction for prediction and are used for predicting the associations between drugs and diseases.The four methods have excellent predictive performance.The final experimental results indicate that more advanced feature extraction methods can significantly improve the predictive performance of the methods,providing new ideas for disease development.

  • 【分类号】R91;TP311.13;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络