节点文献

基因识别算法研究与基因组进化分析

Studies on Gene Identification Algorithms and Analysis of Genome Evolution

【作者】 周立前

【导师】 喻祖国;

【作者基本信息】 湘潭大学 , 应用数学, 2008, 博士

【摘要】 随着人类基因组计划的完成和后基因组时代的到来,生物序列数据呈指数级增长,分析处理大批量数据,从中提取对人类有价值的信息,成为了生物信息学研究的首要任务。我们的工作主要为两个方面:一是区分原核生物完全基因组DNA序列中的编码区与非编码区及人类完全基因组中的基因区与非基因区;二是利用脊椎动物线粒体完全基因组DNA序列与蛋白质序列、多瘤病毒完全基因组DNA序列与蛋白质序列分析物种之间的系统发育关系。本博士论文由四章组成。第一章绪论,主要介绍了生物信息学的概念与研究内容及研究意义、生物信息数据的组成、常用的生物信息处理的数学方法、基因识别算法的概念与当前已有的算法和软件、物种系统发育分析的现状和已有的算法与软件。第二章是关于完全基因组中编码区与非编码区的区分问题,主要综合运用分形、统计、信息等理论和方法,建立处理DNA序列数据的数学模型,应用已有的算法和我们提出的算法分析处理原核生物完全基因组DNA序列和人类完全基因组DNA序列,实现编码区与非编码区、基因区与非基因区的区分。目的在于分析这些基因识别方法的稳定性与高准确率,以期为探索新的未知基因提供新方法、新思想。在原核生物完全基因组的编码区与非编码区的区分中,通过应用了分形方法与Fourier变换方法,获得了较高的区分准确率。在分形方法中,平均区分准确率达78.41%,而Fourier变换方法的区分准确率达到了86.58%。在人类完全基因组的基因区与非基因区的区分中,通过综合应用重分形分析、正四面体、Z曲线和全局描述四种方法,尽管人类完全基因组内部结构非常复杂,仍然获得了高达83.74%的区分准确率。论文的第三章主要介绍系统发育分析的数学模型和方法。第四章应用这些方法去分析处理DNA序列、蛋白质序列等数据集(包括64种脊椎动物线粒体完全基因组序列和70种细菌完全基因组序列),构建物种间的系统发育树,分析各物种间的亲缘与进化关系。在64种脊椎动物线粒体完全基因组和70种多瘤病毒完全基因组的系统发育分析中,我们获得了与传统系统发育树一致的树,综合以前我们的工作发现,我们在系统发育分析研究中提出的方法和模型是可靠的、稳定的,对分析物种间的亲缘与进化关系是非常有意义的。

【Abstract】 Currently, available biologic sequence data are increasing exponentially with the completion of human genomic project (HGP) and the coming of the post genome era. It becomes a very important task in the study of bioinformatics to analyze huge dataset and obtain the valuable information for human. We have worked on two problems. One is distinguishing the coding segments and non-coding segments in the whole genomes of prokaryotes and the gene segments and non-gene segments in the complete human genomes. Another is that analyzes the phylogeny of the vertebrate and polyomaviruses using the DNA and protein sequences of whole mitochondrial genomes and polyomaviruese genomes.The article is made up of four chapters. In the Introduction Chapter 1, the basic conceptions and study content of the bioinformatics and the organization of bioinformatics data are introduced. It also includes the widely used mathematics methods and software of gene identification and species phylogeny in bioinformatics. Chapter 2 is about the problem of distinguishing the coding segments and non-coding segments in the whole genomes. We establish mathematical models to deal with DNA sequence data using the theory and methods of fractals, statistics and information. We apply the existed algorithms and our newly proposed algorithms to distinguish coding and non-coding sequences in the genomes of prokaryotes and human. Our aim is to analyze the stability and high accuracy of these gene finding algorithms. Then try to find some new methods and ideas for gene finding problem. We used fractal method and Fourier method to distinguish the coding and non-coding sequences in the genomes of prokaryotes, the average distinguish accuracy of fractal method can reach 78.41%, while that of Fourier method can reach 86.58%. We also used multifractal (MF), regular tetrahedron (RT), Z curve (ZC) and global descriptor (GD) methods together to distinguish coding and non-coding sequences in human genome. The distinguish accuracy reach 83.74%.In Chapter 3, we introduce the mathematical models and methods to do the phylogenetic analysis. In Chapter 4, we use these methods to study the DNA sequences and protein sequences (include 64 vertebrate mitochondrial genomes and 70 parvovirus genomes), to construct the phylogenetic tree and analyze the evolutionary relationship among species. In the phylogenetic analyses of the 64 vertebrate mitochondrial genomes and 70 parvovirus genomes, we get the phylogenetic trees coincide with those obtained using traditional methods. Hence the phylogenetic models or methods proposed by us are reliable and stable. They are significant for the phylogenetic analysis.

  • 【网络出版投稿人】 湘潭大学
  • 【网络出版年期】2009年 05期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络