节点文献

中国马业综合数据库的建立及马基因组序列预测

Foundation of Horse Synthesis Database of China and Prediction of Horse Genome Sequence

【作者】 乌尼尔夫

【导师】 芒来;

【作者基本信息】 内蒙古农业大学 , 动物遗传育种与繁殖, 2009, 博士

【摘要】 现代生物学的发展促进了生物信息学的产生。生物信息学是将信息学的理论技术应用于生物数据的管理和分析,是数学、物理学、计算机科学、化学、生命科学等多学科的交叉学科。生物信息学研究的范围十分广泛,其中数据库的构建就是一个重要方面。如何用理论和计算的方法识别和预测内含子和外显子也是目前生物信息学研究工作的重要任务。本课题通过自编程序建立了以中国马品种资源为主的中国马业综合数据库www.chinahorse.org.cn。并在建立数据库的基础上,初步实现了数据库应用,包括基于Web的文献数据库的网络化查询等。它将为建立马品种资源的科学研究平台打下基础。本研究的主要内容及结果如下:1.建立了专一化、系统化、完整化的马业科学数据库。序列数据库中以基因数据库和蛋白质数据库为主,非序列数据库以文献数据库和图片数据库为主。其中,马的基因数据库中的记录量超过了2万,马的蛋白质数据库的记录超过3万。2.建立了中国马物种资源数据库。涉及品种的外貌、类型、典型特征等多个性状,为从事中国物种品种遗传资源的利用与保护提供了参考。3.建立了马生物信息学研究平台。可以对基因和蛋白质进行相关生物信息学研究,对于进行科研和教学具有一定价值。4.建立了马业科学实验室网站与马业论坛。可以通过互联网进行数据库的检索,提高了数据库的应用效率。网站的建设还可以为数据库的更新带来方便,也为本研究领域内的交流与合作起到桥梁作用。本研究还通过对已发表的马全基因组序列的密码子使用频率做了初步的统计分析工作并对内含子和外显子进行了预测。基于各种序列组分的不同和序列首尾段的保守性,本研究利用离散增量结合支持向量机的方法对马基因组内含子和外显子序列进行识别。基于单碱基、二联体和三联体使用频率,我们能正确预测91%以上的内含子和外显子。

【Abstract】 With the development of modern biology and a growing accumulation of biological data, a new field of biology-Bioinformatics emerged. Bioinformatics involves applications of the theories and technologies of informatics to biological data analysis and management, and it is a multi-misciplinary research field relating to mathematics, physics, computer science, chemistry and biology. The research of bioinformatics ranges widely and one of the most important fields is database development. And it is an important task for bioinformatics to recoganize and predicts exon and intron using theory and calculation methods.This topic has found the network databases of china horse resources by using programs compiled by us named Horse Synthesis Database of China www.chinahorse.org.cn. Based on network databases, some primary applications of the databases have already been done, including network of bibliographic database based on Web and so forth. These would benefit for the foundation of science research platform of horse breeds resource. The detail information is as follows:1. Specificity, systematic and complete horse science database is established. The sequence libraries mainly include Gene Database and Protein Data Bank. The non- sequence library mainly include Bibliographic Database and Picture Database. The records of Horse Gene Database have over 20,000 and the records of Horse Protein Database have over 30,000.2. The database of Chinese indigenous horse breeds is established. The contents of databases involve their exterior, color and typical characters. The databases would provide some useful information to utilization and protection of Chinese horse breeds.3. Bioinformatacs platform of horse can research genes and their associate features, can simulate protein and there-dimension structure as well. The beautiful output images are easy for educational use and complicated researches.4. Established a horse science and industry of laboratory web and BBS of horse industry. It can be searched by the internet and enhanced the application efficiency of database. Website is a tool of communications, meanwhile, is a platform to update these databases. In this study we did a preliminary statistical analysis on codon usage in the entire horse genome sequence and the intron and exons were predicted. Based on a variety of the different components of the sequence and the conservative character of start and end of the sequece, this study combines the use of increment of diversity and support vector machine method to identify horse genome intron and exon sequences. Based on the single base, diad and triad frequency of use, we are able to correctly predict more than 91% of the intron and exon.

  • 【分类号】S821;TP392
  • 【被引频次】3
  • 【下载频次】399
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络