节点文献
关于生物信息学的几个问题
Several Problems about the Bioinformatics
【作者】 张景祥;
【导师】 徐振源;
【作者基本信息】 江南大学 , 轻工技术与工程, 2008, 硕士
【副题名】DNA序列编码区与非编码区识别方法的研究
【摘要】 随着人类基因组计划的发展,近年来生物信息的数据呈指数增长,如何从大量的数据中挖掘出有用的生物信息是生物信息学领域今后致力解决的问题,其中基因识别问题即通过计算的方法识别DNA序列中编码蛋白质的基因更是十分迫切需要解决的研究课题之一。目前,基因识别常用的方法有:复杂度分析方法、人工神经网络方法、傅立叶分析方法和统计学方法等。概括起来,基因预测方法大致分为两类。一类是基于编码区的碱基组成和非编码区的差异;一类是基于编码区所具有的独特信号,如起始密码子、终止密码子等。本文首先介绍了生物信息学发展情况、基本概念,研究内容和研究方法。然后运用三种寻找CpG岛的方法,找到可能存在基因的位置,并在此基础上,结合一种新的DNA序列字母向量表示方法((?)14),利用信息熵β-KL离散量预测DNA序列的编码和非编码的方法,提高了识别基因编码与非编码区边界的效率,同时,拓展了W-Li阈值的定义,给出S′,通过搜索β=0,0.1,0.2,…,0.9,1,比较发现β∈(0.5,0.7)效果最好。在β=0.65时利用找Dβ-KL找到DNA序列的编码和非编码的边界准确率达到89%,高于Bernalola-Galvan提出的70%的算法,而且计算的时间有显著的减少。
【Abstract】 In recent years,genome projects have given rise to an exponentially growing amount of genetic information.How to find out useful information in the huge amounts of data is the problem that scientists focus on in current and future.One of the most important and basic problems is the gene identification,namely the identification of protein-coding regions in DNA sequences through computational means.In present,a number of methods for gene detection,based on distinctive features of protein coding sequences have been proposed.For example:the method based on correlation function,neural net-based method,Fourier-based analysis,statistics and so on.The methods of the identification of protein-coding regions in DNA sequences can be classified as two kinds,one is based on the difference between the coding and uncoding regions in DNA sequences.Other is based on the signs of the protein-coding regions in DNA sequences,For example:The distribution of the codon and stop-codon.In this study,firstly we introduce the status,basal concepts,researchful content and methods of the Bioinformatics,Then we use the three different methods to find out the CpG island and determine the possible position of the gene.We present a new method to denoting DNA sequences(R14)based on the distributions of "stop-codon" and’reverse complementation stop-codon".Using the theory of Shannon entropy,we ameliorate the measure of the Jensen-Shannon divergence andβ- KL divergence,And Compare with the previous results of experimentation obtained by our method,Showed that recognition efficiency based on the new information measures with the vector(R14) rise 89%,And more than that of by Bernaola’s methods presented 70%.And the time of the calculation is reduced remarkably.
【Key words】 coding and uncoding regions; Bioinformatics; CpG island; (?)14; β-KL Divergence;