节点文献

甘蔗割手密种基因组数据库的构建

SGD:The Sugarcane Saccharum Spontaneum Genome Database

【作者】 陈峥

【导师】 张积森;

【作者基本信息】 福建农林大学 , 生物信息学, 2019, 硕士

【摘要】 甘蔗(Saccharum spp.)是世界上最重要糖料、生物燃料作物,是研究C4光合作用途径和同源多倍体遗传的模式植物,具有巨大的经济和科研价值。在过去的数年中,甘蔗的测序及分析数据迅速积累,为本研究的顺利开展奠定了基础。使用Tripal工具构建甘蔗割手密基因组数据库(http://www.sugarcanetf.site/sgd/html/index.html)作为甘蔗研究的中心门户对这些数据进行存储、挖掘、分析、整合以及共享。研究主要结果如下:(1)甘蔗割手密基因组数据的开发使用BLAST2GO和EggNOG,对甘蔗中99,708个基因进行GO项目注释,65,277个基因进行KEGG生物通路注释。对甘蔗割手密预测基因中的1,278个特异性基因家族进行GO富集,发现这些基因的功能大多富集在对伤口、外部刺激的反应。使用MISA软件对甘蔗割手密进行全基因组SSR开发,共发现577,299个SSR位点,其中染色体特异性位点有98,271个,约占总数的17%。将其与其他四种禾本科植物中的基因组SSR进行比较,发现在禾本科植物中,SSR的丰度与其基因组大小成正比,而SSR的相对丰度与基因组大小没有明显的相关性。开发植物全基因组转录因子预测分类流程,调用HMMER软件实现甘蔗割手密的全基因组转录因子预测及分类。研究中共预测到4,271个编码转录因子的基因,并将其分为57个转录因子家族。(2)甘蔗割手密转录组表达谱数据的开发使用HISAT2和Cufflinks对甘蔗割手密叶段发育、不同生长时期以及昼夜节律的材料进行表达量计算,所得数据可以为甘蔗碳水化合物、光合途径等重要生物学性状基因家族表达谱的研究提供帮助。(3)甘蔗割手密重测序群体基因组数据的开发使用GATK进行变异检测,共识别出448万个高质量的变异型,其中包括约390万个SNPs。之后分别使用SNPhylo和Admixture软件进行系统演化分析和群体结构分析。群体结构分析结果表明可将甘蔗割手密群体分为三个亚群,三个亚群中染色体倍性均呈广泛分布状态。所得数据可用于甘蔗割手密自然群体遗传背景的研究,并为甘蔗育种过程中割手密资源的利用提供帮助。(4)甘蔗割手密基因组数据库的构建基于上述数据集,本研究创建了国际上第一个甘蔗割手密全基因组数据库(Saccharum Genome Database,SGD)。SGD是一个用户友好型的交互式数据库,提供的数据集包括:基因组、蛋白序列、功能注释、表达量、转录因子、分子标记等。除了优质的数据集,SGD还为用户提供了详细的用户手册、强大的搜索工具以及实用的在线工具:JBrowse和BLAST。SGD网站将不断进行数据更新以促进甘蔗及其近缘物种的分子生物学、功能基因组学和遗传进化的研究。

【Abstract】 Sugarcane(Saccharum spp.)is the most important sugar and biofuel crop in the world,and plays an important role for the daily life of people worldwide.Sugarcane is also very important in scientific research community due to special biological characteristic such as the C4photosynthesis,high sugar accumulation,and high biomass.During the past several years,sequencing and genetic data have been rapidly accumulated for sugarcane.In this study,to store,mine,analyze,integrate and disseminate these large-scale datasets and to provide a central portal for the sugarcane research and breeding community,we have developed the Saccharum Genome Database(SGD:http://www.sugarcanetf.site/sgd/html/index.html)using Tripal toolkit.The main results of the study are as follows:(1)The development of the genomic resource in S.spontaneumWe annotated 99,708 genes with GO terms using BLAST2GO and65,277 genes with the KEGG biological pathway using EggNOG in AP85-441.Comparing with the gene families in rice,sorghum,maize,and Arabidopsis,about 1,278 of specific gene families were found in sugarcane.Then we annotated and enriched these genes,the results showed that they were mostly enriched in the response to wounding/external stimuli.In this study,MISA software was used to perform a genome-wide SSR locus search for the S.spontaneum AP85-441.A total of 577,299SSR loci were found,of which 98,271 were chromosome-specific,accounting for 17%of the total.Compared with the number of SSRs in other gramineous plants,we found that the abundance of SSR was positively correlated with genome size,while its relative abundance has no significant correlation with genome sizes.In this study,HMMER was used to predict the transcription factors in S.spontaneum AP85-441,and a total of 4,271 genes for 57 families transcription factors were predicted.(2)The development of the expressional profile for S.spontaneum based on RNA-seqThe transcriptome plays an important role in connecting genomes and proteomes in life science research.In this study,HISAT2 and Cufflinks were used to calculate the transcriptome expression of the SES-208 leaf segment development model,growth and development process,and circadian rhythm(2H).These data can also help the research in the expression profiling of important biological gene families such as photosynthesis and sugar transport in sugarcane.(3)The development of genetic variation based on resequencing of64 S.spontaneum accessionsIn this study,GATK were used to identify variant,and a total of 4.48million high-confidence variants that included 3,961,408 SNPs were mined based on resequencing of 64 64 S.spontaneum accessions.These data provided the resources for the study of the natural group genetic backgrounds and the utilization of breeding parents for sugarcane.(4)Construction of the Saccharum Genome DatabaseBased on the genomic resources we developed herein before described,we established the first S.spontaneum whole genome database(SGD)in the world.SGD is a user-friendly,interactive database that provides datasets including genome,CDS,protein sequences,functional annotations,expression levels,transcription factors,molecular markers.Except for its high-quality datasets,SGD also provides users with detailed user manuals,data integration information and useful online tools including JBrowse and BLAST.The SGD website will be continuously updated to promote the development of molecular biology and genetics of sugarcane and its related species.

  • 【分类号】S566.1;Q811.4
  • 【被引频次】1
  • 【下载频次】343
节点文献中: