节点文献

生物序列的分析方法及其进化模型研究

Research on Analysis and Evolution Model of Biological Sequences

【作者】 解小莉

【导师】 袁志发;

【作者基本信息】 西北农林科技大学 , 动物遗传育种与繁殖, 2012, 博士

【摘要】 自从达尔文时代起,重建地球上所有生命的进化历史,并用系统树的形式来描述这部历史,已经成为许多生物学家的一个梦想。随着分子生物学及生物技术的迅猛发展,生物科学的数据资源急剧膨胀,人们开始利用生物序列数据推断生物进化历史,建立了分子系统学。生物序列及其进化模型是分子系统学的重要组成部分。传统的生物序列分析方法是比对分析方法,20多年前,非比对分析方法作为比对分析方法的补充和发展而出现,并且成为计算分子生物学的一个研究热点。本文以生物序列为研究对象,提出了一些新的分析方法,研究了生物序列进化过程的模型,为进化距离的改进提出了一些新的结果。本文的主要工作包括以下几个方面:1.提出了新的氨基酸序列的两种图形表示方法。一种方法是按照疏水性把氨基酸分为三类,对分类后的氨基酸序列进行分析,文中定义了三条曲线(IA曲线,EA曲线和IE曲线),这三条曲线不仅能可视化新的序列,而且可以比较三类氨基酸的分布情况。首次引入条件概率作为序列的数字特征,以挖掘隐藏在序列中的相关性信息,并在此基础上分析了序列的相似性。第二种图形表示方法是直接利用氨基酸的疏水值,定义了氨基酸序列的疏水性曲线,根据传统的图形量化方法分析序列的相似性。以ND6(NADHdehydrogenase subunit6)蛋白质序列为生物数据,阐明并比较了两种分析方法,结果表明两种方法是合理可行的。2.提出了蛋白质序列的选择进化距离公式。蛋白质序列的进化距离因氨基酸替代模型的不同而不同,针对这一问题,本文在选择的思想下,建立了氨基酸序列的进化模型,得到了新的进化距离(选择进化距离)公式,给出了进化距离中参数的确定方法。通过17个物种的细胞色素b的氨基酸序列说明了选择进化距离的计算方法,并根据自展法比较了不同进化距离得到的物种进化树。结果表明,利用选择进化距离构建的进化树与其他几种进化距离得到的进化树的拓扑结构一致,而且选择进化距离的计算避免了氨基酸替代模型的选择问题。3.提出了量化DNA序列碱基分布的指标。在DNA序列的4个单碱基和16个双碱基的距离分布模型的基础上,定义了4个单碱基和16个双碱基的平均距离指标和相对距离熵指标,以衡量DNA序列中所有单(双)碱基的分布情况。平均距离指标描述了相邻两个相同单(双)碱基之间其他碱基的数目,相对距离熵指标描述了每个单(双)碱基的距离分布的均匀性程度。在此基础上,分析了17个物种的线粒体基因组序列的单(双)碱基的分布情况,在线粒体基因组序列中,相邻两个碱基G的平均距离和相对距离熵比其它碱基的平均距离和相对距离熵大;双碱基CG的平均距离比其他双碱基的平均距离大。4.建立了核苷酸替代的转换-颠换进化动力学模型。在核苷酸替代动力学模型及研究结果的基础上,分析了不同核苷酸替代模型下核苷酸序列的相同对、转换对和颠换对频率随时间的变化规律,说明了转换颠换比和各种进化距离的特征,建立了核苷酸序列的转换-颠换进化模型,给出了模型中各个参数的确定方法及其生物学意义。和传统的核苷酸替代动力学模型相比,该模型优点在于更加简单,由核苷酸替代阵或核苷酸替代路径图可以直接得到,而且这个模型很容易给出DNA序列的进化距离。

【Abstract】 Beginning from Darwin’s era, it has become a dream for many biologists to reconstructthe evolution history of all the species in the world and describe the history by usingphylogenetic tree. The data resources of biological science are rapidly expanding with thedevelopment of molecular biology and biological technology. People began to detect theevolution history of living beings by biological data instead of phenotype, and built molecularsystematics. Biological sequence and sequence evolution models are the key components ofmolecular systematics. The traditional method for biological analysis is sequence alignment,as its complement, over20years ago, the free-alignment method emerged, which has becomea hot issue of computational molecular biology. This dissertation chose biological sequencesas a research question, put forward some new method to analyze biological sequence, andstudied the model of biological sequences evolution process to provide some new results forimprovement evolution distance. The main content includes several aspects as follows:1.New graphical representations of protein sequence were put forward. In the firstgraphical representation method,20amino acids were divided into three groups according totheir hydropathy and the new amino acid sequence was analyzed. Three curves were defined,namely, IA curve, EA curve and IE curve. The three curves can not only make new sequencevisible, but also compare the distribution of three types of amino acids. To quantify thesequence, we first introduced conditional probability as numerical characterization forsequence so as to probe into the related information hidden in the sequence. Thesimilarity/dissimilarity between protein sequences were analyzed by basing on conditionalprobability. The second method directly utilized the hydropathy of amino acids and definedthe hydropathy curve of amino acid sequence, and analyzed the similarity between differentsequences by employing traditional diagram quantitative analysis. Using ND6(NADHdehydrogenase subunit6) protein sequence as biological data, the paper carried comparativeanalysis on the two methods, which indicated that the two methods are feasible and effective.2.The selection evolution distance formula of protein sequence was proposed. Under thetheme of selection,the dynamic equation of amino acid sequences was built, and the selectionevolution distance was obtained. The parameter of the model was given. The amino acid sequences of cytochrome b in17species were taken as an example to illustrate the newevolution distance, and the different evolution trees based on different evolution distanceswere compared by the bootstrap. The result showed that the topology of the evolution treebased on selection evolution distance was consistent with that based on other distances, andthe estimation of selection evolution distance avoided the choice among different amino acidsubstitution models.3.The indexes to quantify nucleotide distribution of DNA sequences were proposed. Thedistance distribution models of4single nucleotides and16dual nucleotides in the DNAsequence were provided. Based on this, the average distance and relative distance entropy ofthem were also put forward to test the distribution of each single/dual nucleotide in DNAsequence. The average distance described the number of other nucleotides between theneighboring two single/dual nucleotide, while the relative distance entropy described thedegree of evenness for single/dual nucleotide’s distance distribution. As an application, weanalyzed the nucleotide distribution of17species’ mitochondrial genome sequences. Inmitochondrial genome sequences, the neighboring two nucleotide G’s average distanceentropy and relative distance entropy were greater than that of other nucleotides, and dualnucleotide CG’s average distance was larger than that of other dual nucleotides.4.The conversion-transversion evolutionary model of DNA sequences was built. On thebasis of nucleotide substitution dynamic model, this paper analyzed the pattern of changes forfrequency of nucleotide sequences’ similar pair, conversion pair and transversion pair withchanges of the time. The characteristics of transversion-conversion ratio and evolutionarydistances were also analyzed. The conversion-transversion evolutionary model was putforward. The method for estimating parameters in the model and its biological significancewere also provided. Compared with the traditional dynamical model of nucleotide substitution,the establishment of this model is simpler, which can be obtained directly by the nucleotidesubstitution matrix or nucleotide substitution path map. Through this model, the DNAsequence’s evolutionary distance was easily obtained.

节点文献中: