节点文献

基于逻辑回归的药物与靶标相互作用关系预测

Drug Target Interaction Prediction Based on Logistic Regression

【作者】 张丽娟

【导师】 陈湘涛;

【作者基本信息】 湖南大学 , 计算机科学与技术, 2018, 硕士

【摘要】 药物发现和药物重定向是一项花销很高并且很耗时的进程,而推进药物发现和药物重定向进程发展的一个有效的方法是从分子层面上对药物和靶标的相互作用关系进行预测。传统的方法对药物和靶标的相互作用关系进行预测存在特征难以获取以及计算复杂度高等问题。因此,近年来,许多统计学习方法被用来解决药物靶标相互作用关系预测问题,尽管他们在预测上很有效,但是它们仍存在以下问题:首先是缺少负样本的问题,在现在所有的记录药物靶标相互作用关系的数据库中,均只有已经经过实验验证的药物靶标相互作用关系数据,而经过实验验证的药物靶标之间不存在相互作用的数据在任何数据库中都没有记录。其次是药物靶标的特征选取以及表示问题,主要是考虑如何选择合适的特征以贴合的表达一个药物一个靶标,怎样组合药物特征和靶标特征才能完整表达一个药物靶标对。为了解决上述的两个问题,本文做了如下工作。1)在药物靶标对特征构建的过程中,本文针对药物结构特征矩阵和靶标结构特征矩阵它们内部存在相互依赖的问题,提出了一种适用于0-1矩阵的信息增益主成分分析法(Information Gain Principal Component Analysis,IGPCA),并提出了三种基于药物靶标相互关系矩阵构建网络特征的方法。最后,本文采用五折交叉验证法对该方法进行了有效性验证,实验结果表明,使用IGPCA处理后的结构特征在各种分类器下能取得很好的效果,网络特征在大部分分类器下是有效的。2)为了解决负样本缺失以及特征表达的问题,本文提出了二维逻辑回归模型(Two Dimensional Logistic Regression Model,TwoDLLR)用于药物靶标相互作用关系的预测。该模型将二维计算模型引入逻辑回归方程,并利用Pearson相关系数和随机梯度下降法求解模型参数。模型能够很好的适应药物靶标关系预测问题的药物特征和靶标特征分开的情况,并且该模型能够不使用负样本,解决了药物靶标相互作用关系预测问题中的不存在负样本的问题。五折交叉验证的实验结果表明,本文提出的模型在AUC、Accuracy和F-score指标下都能取得很好的效果,并且TwoDLLR模型能够预测新药物的相互作用靶标以及预测新靶标的相互作用药物。

【Abstract】 Drug discovery and drug reposition are time-consuming and expensive process all the time,and identifying drug-target interactions that considered from the molecular level is an effective method for drug discovery and drug reposition.However,finding interactions between drugs and targets is difficult for the difficult of feature design and the complexity of computing.Recently,many statistical models effectively inferred novel drug-target interactions,but they suffered some limitations.Firstly,data sets used only contain true-positive samples,and experimentally validated negative samples are not available,because database about drug target interactions in internet only contain the certain interactions of drugs and targets,while the opposite interactions between drug and target does’ t exist in database.Secondly,the problems of expressing and selecting drug features and target features.It is mainly to consider how to select the appropriate features and express a drug and a target,and how to combine features of drug target pair.In this thesis,we discussed feature processing of drugs and targets,and we proposed a concept of network features.To predict interactions of drugs and targets and solve the problems of samples without negative samples,we improved a linear logistic regression and proposed two dimensional logistic regression model.the main work described as follows:1)In order to express the feature of drug target pair more accurately,we put forward the concept of network features.First,we obtain the structure features of drugs and targets.We use information gain principal component analysis(IGPCA)for drug structure features and target structure features to perform independence.Then,taking into account the information of drug target interaction,we extract drug network features and target network features from the matrix of drug target interaction.Finally,we verify the effectiveness of the feature processing method by several group of five-fold cross validation.The experimental results show that the principal component analysis method can achieve good results under various classifiers,and the extracted network features are effective under most classifiers.2)In order to solve the problem of negative sample and feature expression problems,we propose a new model——two dimensional logistic regression model(TwoDLLR),which improves the traditional logistic regression model.It can be well adapted to the case of the drug target relationship prediction,and the model can predict interactions without negative samples.Finally,we verify the effectiveness of our method by several group of 5-fold cross validation experiments.The experimental results show that our method can achieve good results under different classification evaluation index,and results show that TwoDLLR method is able to discover most of interactions of new drug and interactions of new target.

  • 【网络出版投稿人】 湖南大学
  • 【网络出版年期】2019年 02期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络