节点文献

基于智能计算的蛋白质功能预测研究

Study on Protein Function Prediction Based on Intelligent Computation

【作者】 张同亮

【导师】 丁永生;

【作者基本信息】 东华大学 , 控制理论与控制工程, 2008, 博士

【摘要】 蛋白质是生命体赖以生存的营养要素,是细胞组织的重要组成部分。几乎所有的生物过程都与蛋白质发生某种联系。根据蛋白质序列的排列顺序和序列信息确定蛋白质的功能成为生物学研究重点。目前蛋白质序列数量的激增,急需要开发快速、准确地计算工具预测蛋白质的功能。研究蛋白质序列信息与其功能的关系也是这个领域的研究重点。本论文围绕蛋白质功能预测的几个重要方面:蛋白质亚细胞位点预测,蛋白质结构类预测和单序列蛋白质二级结构预测和蛋白质序列内功能Motif发现展开研究,目的是开发一些根据序列信息预测蛋白质功能的方法。论文的主要研究成果如下:在蛋白质亚细胞位点预测研究中,根据Chou提出的伪氨基酸组成离散模型,提出一种改进的伪氨基酸组成模型。使用免疫遗传算法优化附加特征向量的权重。在改进的伪氨基酸组成模型框架中,使用数字信号处理技术和疏水氨基酸对模式表示序列的附加特征,应用扩大的协方差作为预测工具,预测了真核细胞12类亚细胞位点。然后提出了一种基于特征选择的集成分类器的预测方法,用于凋谢蛋白的亚细胞位点预测。使用具有不同间隔的氨基酸对组成表示序列特征,经过特征选择后形成更加有效的特征组合。集成分类器中的基本分类器为模糊K-近邻(FKNN)分类算法,Jackknife测试和独立数据集测试证明了该方法的有效性和实用性。在蛋白质结构类预测研究中,提出了三种结构类预测的方法。第一种是基于二叉树支持向量机的方法,发展了一种新的伪氨基酸组成表示序列的特征。结合了传统的氨基酸组成,序列内氨基酸相互关系和疏水模式,使用二叉树支持向量机作为预测工具,采用标准数据集验证了方法的性能;第二种方法是基于改进的伪氨基酸组成模型的结构类预测方法。将蛋白质序列映射为短的时间序列,计算序列的近似熵,构造了一种27-D的伪氨基酸组成表示序列特征。FKNN分类算法作为预测工具,免疫遗传算法优化附加特征权重系数。在“严格”数据集测试中取得了较好的结果;第三种方法是两层模糊支持向量机网络的方法,在第一层中,基本的分类器是模糊支持向量机,输入数据是基于不同物理化学属性的伪氨基酸组成。组合第一层中各个模糊支持向量机的输出数据,作为第二层模糊支持向量机分类器的输入数据,经过决策后得到最终结果。在蛋白质二级结构预测研究中,提出了基于最大熵概率模型的预测方法。考虑了蛋白质序列的结构类信息和目标残基的上下文环境,设计了影响残基二级结构的特征空间和特征模版。将这些特征都包含进入最大熵概率分布模型中,根据结构类不同分别训练和建立二级结构预测模型。算法中二级结构的特征信息仅来自于序列本身,没有考虑多序列排列信息。目的是解决“孤立”蛋白的二级结构预测问题。实验证明预测算法具有较高的准确率和实用性。由于细胞核内空间狭窄和蛋白质的不稳定性,核内亚空间的蛋白质位点预测成为难点。本论文提出了基于近似熵的伪氨基酸组成方法,采用集成AdaBoost分类器作为预测工具,用于蛋白质亚核位点的预测。在两个标准数据集上的测试表明了该方法的有效性。蛋白质家族内序列具有相似的功能,序列内的重点区域Motif也应该具有相似性。本论文提出了一种Motif发现算法,在蛋白质家族内寻找重要的Motif集合,验证序列所属的蛋白质家族。在连接酶的21个亚家族识别中,建立了一个实用的连接酶亚家族服务器。最后,对全论文的研究内容进行了总结,指出了研究工作中存在的不足,明确了下一步的研究方向。

【Abstract】 Protein, which is an important part in cell, plays a critical role in life processing. Protein function prediction, i.e. classification of protein sequences according to their biological function is an important task in bioinformatics and protein science. The gap between the numbers of known protein sequences and the number of annotated protein is increasing rapidly. It is highly desired to develop some powerful tools and effectively methods to bridge the gap. Prediction of protein function with computational approaches is one of the most important research topics in protein science and bioinformatics. Meanwhile, finding the knowledge of relationship between protein sequence and its function is an important research field. This thesis mainly focuses on several important problems in prediction of protein function: protein subcellular localization, protein structural classes, and protein secondary structure prediction. We aim to develop some approaches to predict protein function from its sequence. The main contributions in the thesis are described as follows.First, we investigate the development of protein subcellular localization prediction.Its difficulties and further developments are summarized. According to the concept ofPseudo Amino Acid (PseAA) composition originally introduced by Chou, we proposean approach of improved PseAA (IPseAA) composition in which the weight factors areoptimized by immune genetic algorithm. Based on the approach of IPseAA, a novelfeature vector is developed to represent the sample of protein which incorporates theconcept of average power-spectral density and hydrobolicity pattern. Promising resultsare obtained when the method is used to predict eukaryotic protein subcellularlocalization. Then, we propose another approach to predict apoptosis proteinsubcellular localization. An ensemble classifier is proposed, in which the basicclassifier is fuzzy K nearest neighbors (FKNN) algorithm. Each basic classifier istrained by collocated amino acid pair composition. The collocated amino acid pair is apair amino acid with different spaces. Feature selection algorithm based on geneticalgorithm is used to get the optimazed features. The results of Jackknife and independent dataset tests indicate that the proposed approach is effective and practical.For prediction of protein structural classes, we propose three methods for it. 1) Based on binary-tree support vector machine (BT-SVM). Combined amino acid composition, correlation of amino acids in sequence, and hydrobolicity pattern, a novel PseAA composition is developed to represent sample of proteins. BT-SVM is used as prediction engine, which has capability in solving the problem of unclassifiable data points in multi-class SVMs. 2) Based on the concept of the approximate entropy (ApEn) and hydrophobicity patterns a novel approach is proposed to generate the PseAA composition for protein samples. FKNN classifier is used as prediction engine. A large and stringent dataset is adopt to validate the performance of the approach, encouraging results indicate the novel PseAA composition based on the concepts of ApEn and hydrophobicity patterns might reflect the core feature of proteins in different structural classes. 3) A two layers fuzzy support vector machine (FSVM) network is proposed to predict protein structural classes. In the first layer, the input data of the basic classifier (FSVM) is the PseAA composition based on different physi-chemical properties of amino acid. The outputs of FSVM in the first layer are combined into a vector. It is the input of the FSVM in the second layer.Nature language processing methods are introduced to handle the problem of protein seconday structure prediction. We propose an approach to predict protein secondary structure based on maximum entropy model. According to the contextual information of target residue and structural classes’ information of protein sequence, feature space and feature templates are designed. All features, which are combined into an event, are incoporated into maximum entropy model. The models trained by the datasets with different structural classes, respectively. The features of protein secondary structure do not use any information from multi-profile, and the aim of the study is to help improving the function annotation of "orphan" protein which has no detectable homologs. Validated by the benchmark datasets, high predictive success rate denotes the approach might become a useful tool in related area.There are few studies on protein subnuclear localization prediction beacuse the nuclear is more compact and complicated as compared to other cell compartments. We develop an approach of ensemble of AdaBoost classifier. The PseAA based ApEn of sequence is used to represent the features of protein sequence. Two benchmarks are used to validate the performance of the approach. Compared with published works, the highest accuracy is achieved.The protein sequences in same family have same function. We can assume that some similarly important regions, Motifs, existing in sequences belong to same family. A Motif discovery method is proposed. In a protein family, a Motif set is searched to reprensent the family. The method has been used to identify ligases 21 subfamly. A Web-server is released for free science study.At last, a summary of the thesis is made, and the deficiency in the project and the further development are narrated respectively.

  • 【网络出版投稿人】 东华大学
  • 【网络出版年期】2008年 12期
节点文献中: