节点文献

基于EST全基因组定位的基因结构注释研究

Research on Gene Structure Annotation Using Genomically Aligned EST Sequences

【作者】 井辉

【导师】 周艳红;

【作者基本信息】 华中科技大学 , 生物信息技术, 2007, 硕士

【摘要】 识别蛋白编码基因是基因组研究中的重要课题之一。特别是随着越来越多的物种被测序,这一课题更加重要。面对急剧膨胀的基因组序列,传统生物学实验已经无法满足需要。因此,生物信息学的高通量方法显得尤其重要。EST(序列表达标签)是对随机选取的cDNA克隆进行测序的一部分,理论上EST不含内含子,代表了一个完整基因的一部分。EST数据量巨大且还在迅速增长之中,是一种宝贵的序列资源。利用EST对基因组进行蛋白编码基因的预测和注释是重要的研究课题。但EST序列的质量问题和基因组序列的复杂性使得这一工作并不容易开展。本研究首先了解了EST序列的产生过程和序列特点,深入分析了可能影响EST序列质量的因素。包括外源序列、基因组DNA序列、嵌合EST序列,mRNA前体序列、随机引导序列、内部引导序列等等。同时对基因组序列也进行了深入分析,包括重复序列成份、假基因、多拷贝基因、重叠和嵌合基因、选择性剪接等等。在此基础上,本研究考虑了EST与整个基因组进行序列比对和定位可能产生的情况,针对这些情况制订了对策和研究方案,具体是:先对EST去除外源污染,然后将其定位到基因组上,并对比对结果采取针对性的措施加以检验;对保留下来的EST,根据相互之间的联系进行聚类,最后预测出基因结构,并利用有向无环图(Directed Acyclic Graph)和期望最大值算法(Expectation-Maximization)得到可能的选择性剪接。本研究取得了令人满意的结果,测试表明,研究中制订的措施是有效的。本研究还设计了一个覆盖整个基因组的基因注释系统,建立了一个包含有约6000万条目的数据库,支撑相关的web服务(http://bioinfo.hust.edu.cn)。

【Abstract】 Identification of protein coding genes is a crucial issue of genome research. As genomes of more and more species have been sequenced, the issue is becoming particularly important. Traditional biological experiments could hardly tackle the whole problem with the explosion of genomic sequences, which makes the high throughput methods of bioinformatics invaluable.An EST (expressed sequence tag) is a partial sequence of a clone picked at random for cDNA library. In theory, an EST contains no introns and represents part of a gene. The amount of ESTs is very large and is growing fast. It is an extremely precious resource. Prediction and annotation of protein coding genes using genomically aligned ESTs is a crucial issue. However, this is not a trifle due to the poor quality of EST sequences and the complexity of genome.Characteristics of EST sequences have been studied at first. Several factors concerning the quality of EST sequences have been investigated, including foreign sequences, genomic DNA sequences, chimeric sequences, pre-mRNA sequences, random-primed sequences, internal-primed sequences and so on. Components of the genome have also been studied, such as repeats, pseudo genes, multi-copy genes, overlapping genes, nested genes and alternative splicing.Based on these analyses, several measures have been brought up according to different cases involved in the alignments between EST sequences and the genome. The annotation procedures include trimming foreign sequences, locating ESTs on the genome, verifying and testing the alignments, clustering between each other and predicting the final gene structure. At the prediction step, directed acyclic graph (DAG) algorithm and expectation-maximization (EM) algorithm were applied to predict and sort out alternative spliced transcripts according to their calculated probability. Evaluation of the results proves that the procedures are effective. The gene annotation platform that covers the whole human genome has been set up. An established large database which contains more than 60 million entries is now supporting related web services at http://bioinfo.hust.edu.cn.

【关键词】 EST序列比对基因组基因预测基因注释
【Key words】 ESTsequence alignmentgenomegene predictiongene annotation
节点文献中: