节点文献

从头预测氨基酸相互作用及其在蛋白质三维结构建模中的应用研究

Ab-initio Prediction of Residue Contacts and Its Application to Protein3D Structure Modeling

【作者】 杨静

【导师】 沈红斌;

【作者基本信息】 上海交通大学 , 控制工程, 2014, 硕士

【摘要】 蛋白质作为生命活动的物质基础,在生物体的各种细胞过程中起着十分重要的作用。研究显示,蛋白质的功能和其空间结构是紧密相联的。因此,如果解出了蛋白质的三维结构,将有助于了解它的生物功能,然后进一步指导药物设计。近些年来,尽管通过实验方法解出结构的数量在不断地增加,但是已知蛋白质序列的数量和蛋白质结构的数量之间的差距也在不停地变大。幸运的是,随着机器学习和数据挖掘技术的快速发展,利用计算的方法直接从蛋白质序列预测出其结构已成为一种流行的方法。在蛋白质结构预测领域内,现有的方法通常可以划分为三类:同源建模的方法、折叠识别的方法和从头预测的方法。对于第一类和第三类方法,都有一个挑战性的问题,就是如何得到蛋白质序列中氨基酸之间的关联信息,而这些信息将被用作蛋白质三维结构建模的限制条件。在同源建模的方法中,氨基酸相互作用信息大多来源于蛋白质结构数据库(PDB)中的同源结构;而在从头预测的方法中,氨基酸相互作用信息主要来源于基于序列的预测结果。显然,当没有找到合适的同源结构时,采用计算方法预测出来的结果将更加精确。在最近的几十年里,人们提出了很多种预测氨基酸相互作用的方法。然而,预测精度并不能令人满意。本文从蛋白质序列出发,利用机器学习的技术来预测氨基酸相互作用,并结合序列比对的方法实现了相互作用的准确预测,再将预测到的相互作用信息作为限制条件进行蛋白质三维结构建模。实验结果表明,预测到的相互作用信息在三维结构建模中是有效的,也因此提高了蛋白质结构预测的精确度。具体来说,本文解决了蛋白质结构预测中的两个问题:(1)跨膜螺旋(TMH)之间氨基酸相互作用的预测;(2)二硫键连接模式的预测。对于TMH之间氨基酸相互作用的预测,本文提出了一种融合机器学习和共变异分析的新方法。这里,我们用偏相关性分析来计算共变异分数,而机器学习模块是由集成分类器来实现。在此项工作中,共变异分数首次在决策层融合。实验证明,这两种方法具有高度互补性,因而大幅提高了预测精度,比现有最好的方法要高12.5%。对于二硫键连接模式的预测,在已知半胱氨酸是否参与形成二硫键的条件下,本文设计出了一个新的融合模型。最终的结果是由机器学习的预测结果和序列比对的分配结果共同决定的。通过引入结构特征,可以提高传统机器学习模型的预测性能。此外,我们首次提出了基于序列比对的方法,直接根据标注过的同源序列进行二硫键的分配。二硫键的连接模式取所有可能的连接模式中概率最大的。实验表明,序列比对的方法可以辅助机器学习的方法,使得最终的预测模型具有鲁棒性,并能够取得很高的预测精度。

【Abstract】 Proteins play essential and important roles in various crucial cellularprocesses in any living organism. It has been revealed that the proteinfunction is closely related to its structure. Therefore, knowing proteinstructures can help to understand their functions and then guide drugdesign. In the recent years, although the increasing number ofexperimentally solved structures, the gap between the available numberof protein sequences and known structures continues to increase.Fortunately, with the rapid advance in machine learning and data miningtechnologies, it is feasible to predict protein structures from proteinsequences directly.In the literature, the proposed methods for protein structureprediction can be generally grouped into three categories of homologymodeling, fold recognition, and ab initio prediction. For the first and lastclasses of approaches, one common challenging problem is how togenerate the residue-residue contact map that will be further used asconstraints in protein structure assembly. In the case of homologymodeling, residue contact information is mainly derived from the knownhomologous structures in the protein data bank (PDB); while for the abinitio prediction, the contact map is mostly obtained from thesequence-based predictions. Obviously, ab initio predictions are moreaccurate when homologous structures are not available.In the last decades, many approaches have been proposed forresidue-residue contact prediction. However, the prediction accuracy isfar from satisfaction. In this paper, we predict residue contacts from theprimary sequence by merging machine learning method with sequence alignment approach and then use them as constraints for protein structuremodeling. Experimental results indicate that the predicted contacts arevalid for structure modeling and hence can improve the accuracy ofprotein structure prediction. Concretely, this work is consisted of twotopics:(1) inter-transmembrane helix (TMH) residue contact prediction;(2) disulfide bond connectivity pattern prediction.For inter-TMH residue contact prediction, we present a new methodthat merging machine learning-based method with correlated mutationanalysis-based approach. Here, we use the partial correlation analysis tocalculate correlated mutation score. The machine learning-based enginein the proposed protocol is implemented with ensemble classifier. It is thefirst time that correlated mutation score is fused in decision level. Theresults demonstrate that these two engines are highly complementary toeach other and hence improve the prediction accuracy, which is12.5%higher than the current best method from the literature.For disulfide bond connectivity pattern prediction, we present anovel consensus model to predict disulfide bonds with known bondingstates of cysteines. It is the fusion of machine learning-based predictionsand sequence alignment-based annotations. We improve the traditionalmachine learning-based model by introducing the feature of structuraldistance information. In addition, we firstly propose a baseline predictorbased on sequence alignment to assist machine learning-based method.The disulfide bond connectivity pattern is predicted by maximizing thesum of probabilities of possible disulfide bonds. The results show that thecombination of these two methods drives the final robust model achievinghigh prediction accuracy.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络