节点文献

茶树全器官转录图谱和基于WGS的全基因组特征的初步研究

Transcriptome Profiling of Whole Organs and Preliminary Study on the Characterizations of Whole Genome of the Tea Plant (Camellia Sinensis) by WGS Technology

【作者】 杨华

【导师】 宛晓春; 韦朝领;

【作者基本信息】 安徽农业大学 , 茶学, 2015, 博士

【摘要】 茶树(Camellia sinensis)是世界三大饮料作物之一,起源于中国。茶具重要的经济价值和世界影响力,而且富含儿茶素、咖啡碱和茶氨酸等特征性次生代谢物,这不仅决定了茶的滋味和品质,而且是茶具有重要保健作用的物质基础。由于茶树自交不亲和,基因组高度杂合且较大,茶树基因组研究基础比较薄弱,缺少全基因组序列图谱;茶树在全基因组广度的转录图谱也不够完善,且遗传转化体系尚未完全建立,这些因素极大地限制了茶树生物学的研究,目前不仅克隆和验证的功能基因相当有限,茶树特征性次生代谢物产生和调控的分子机制也尚未完全解析。本研究旨在采用近年发展起来的下一代高通量测序技术在国内外首次构建茶树基因组框架图和全器官转录组图谱,并进行茶树和野生近缘种基因组广度的SNP发掘。茶树基因组和转录组数据库的建立不仅使得解析茶树基因组特征和进化成为可能,还将对茶树生物学研究(尤其是次生代谢形成和调控机制的研究)提供前提条件和重要平台。本论文的主要研究结果如下:1.茶树全器官转录组深度测序揭示茶树特征性次生代谢途径相关基因本研究利用RNA-Seq技术在国内外首次建立茶树全器官转录组。通过Illumina测序平台,对栽培茶树国家级良种龙井43的全组织器官混合样本进行深度测序,测序数据量达到2.59Gb,经从头(de novo)拼接后得到127,094条Unigene,平均长度355bp,N50长度为506bp。利用GenBank中的茶树EST不仅评估了茶树转录组数据质量的可靠性,还评估了数据量的高覆盖度。利用七个公共数据库,55,088条Unigene获得基因注释、GO(Gene Ontology)功能注释或KEGG代谢途径注释。发现了很多参与茶树初级代谢以及类黄酮、茶氨酸和咖啡碱次级代谢途径涉及的主要基因,并发现一些参与次级代谢的新的候补基因。对茶氨酸和类黄酮合成途径上相关的13种基因进行实时定量PCR分析这些基因在茶树不同组织中的表达模式。此外,基于本转录组Unigene搜索出12,242个SSR,包括单核苷酸重复到六核苷酸重复的各种类型,以二核苷酸重复类型为最多,占所有SSR的63.78%。此外,还对茶树转录组SSR的多态性和可用性进行了分析评价。2.茶树和4个野生近缘种(变种)的基因组SNP发掘及高分辨的物种鉴定和遗传关系研究本研究利用酶切位点相关DNA测序(RAD-Seq)技术对18株栽培茶树和野生古茶树(野生近缘种茶树个体)进行了高通量的简化基因组测序,每个样本获得的有效数据量平均达到2.94±0.67 Gb。对测序标签聚类和核苷酸位点的基因分型分析后获得15,444个高置信度的二等位SNP,并以此对18个供试样本进行系统发生关系分析、主成分分析和群体遗传结构分析,结果比较一致,表明基因组SNP能对供试样本中茶种和4个野生近缘种(变种)进行高分辨的物种鉴定,并了解其相互的亲缘关系。本研究首次在基因组水平上提供分子证据,推测大理茶邦崴变种很可能来自阿萨姆种和大理茶种的种间杂交。除了大理茶邦崴变种,供试栽培型茶树基因组杂合度普遍高于野生古茶树。此外,在获得的15,444个SNP位点中鉴定了1,521个基因区SNP,其相关的基因中1,058条被注释为拟南芥同源蛋白,24条涉及次生代谢过程。3.茶树基因组框架图绘制及基因组特征和进化的初步分析本研究在绘制茶树全基因组框架图前,先对栽培茶树品种安徽1号(AH1)、铁观音(TGY)和舒茶早(SCZ)和1株大理茶野生茶树(DXS)进行基因组调查测序。AH1、TGY、SCZ和DXS获得的有效调查测序数据分别为105.3、110.8、205.1和159.8Gb;采用17-mer估算4者基因组大小约在3-3.3Gb;4者的基因组杂合率约在1.5-2.2%,其大小排序为AH1>TGY>SCZ>DXS;4者的基因组GC含量分别约为38.50-39.92%;4株茶树基因组初步组装的contig N50长度均<700bp,scaffold N50长度均<3kb。在AH1和TGY初步组装序列中分别预测得到563,680和545,520个基因组SSR(gSSR)位点,并批量设计了262,807和257,564对引物。选择杂合度相对较低的栽培茶树SCZ为研究材料,结合全基因组鸟枪法策略和高深度测序和针对复杂基因组组装策略进行茶树全基因组测序研究。构建170~800bp短插入片段文库和2~20kb长插入片段文库,进行Illumina双末端测序,获得有效数据1,393 Gb。同时,对SCZ的8个组织器官进行转录组测序,获得有效数据量94.3Gb。对茶树全基因组测序数据进行de novo组装,产生contig的N50长度为33.4kb,contig总长度为2.45Gb;最终组装了92,207条scaffold,N50长度为347.1kb,总长度为2.98Gb。采用Genbank中登录的茶树EST评估组装的基因区覆盖度可达到约89.25%。随机的挑选3条茶树BAC序列(单独采用Sanger测序)和基因组组装的一致性较高,覆盖率分别为97.7%、100%和94.8%。利用RAD-Seq产生SCZ标签序列评估组装的基因组限制性酶切位点区域,覆盖率达到96.79%。茶树基因组重复序列约占基因组的56.06%,大部分为TE(占基因组52.69%),主要类型为LTR。预测茶树蛋白编码基因48,682条,平均基因长度为4,054bp,每个基因平均3.3个外显子,其中39,680条基因获得注释。此外,预测了652条tRNA、3,450条rRNA、723条miRNA和474条snRNA。基因家族分析表明33,922条茶树基因聚类到17,701个基因家族中,茶树特有的基因家族达到3,415个,茶树、猕猴桃、无油樟和桉树共有基因家族7,171个。茶树基因组进化分析表明表明茶树属于菊类分支,并且与猕猴桃科的亲缘关系最近,并估算两者的分化大约在近期的67.5百万年前发生。基于茶树基因组、转录组和BAC文库,鉴定和克隆了儿茶素合成关键基因LAR的基因组DNA序列,并分析了其基因结构。

【Abstract】 Tea is one of the most popular three traditional beverages worldwide.The ancestors of cultivated tea plants(Camellia sinensis)are native to Southwest China.At the present,tea plants are cultivated in more than fifty countries,and 2 billions people(1/3 of all)drink tea everyday in the world.Besides important economic value and influence in the world,aboudant secondary metabolits,such as catechins,caffeine,theanine and volatile oils,exist in tea,that not only play a crucial role in tea quality and flavour,but also are the essential material basses for promoting human health.However,C.sinensis possess large genomes and high heterozygosity.In addition,the tea plant is difficult to culture in vitro and to transform,which tremendously hinder the research on the genetic engneering of functional genes.To date,the lack of genomic information and the genome-wide transcription profiling imposes large restrictions on biology and molecular genetics studies,especially the biosynthesis of tea-apecific secondary metabolites and genetic regulation mecahnisms in tea plant.This study was designed to construct the first C.sinensis draft genome and transcriptome profiling of all tissues and a large-scale development of genome-wide SNPs from C.sinensis and its several wild relatives in section Thea of genus Camellia.C.sinensis draft genome,transcriptome dataset and derived SSR and SNP resource can serve as an important public information platform for biological characteristics,origin and evolution,functional genomic studies and molecular breeding in C.sinensis.It will tremendouly promote not only the secondary metabolism research in C.sinensis but also the entire development of tea production.The main results in this study were simpliy described as follows:1.Deep sequencing of the Camellia sinensis transcriptome revealed candidate genes for major metabolic pathways of tea-specific compounds and identified SSRsUsing high-throughput Illumina RNA-seq,the first transcriptome profioling of C.sinensis was constructed in the world.Deep sequencing from poly(A)~+RNA of all tissues of C.sinensis cv.Longjing43 was analyzed at an unprecedented depth using Illumina sequencing platform.Approximate 2.59 gigabase pairs(Gb)of reads were obtained,trimmed,and assembled into 127,094 unigenes,with an average length of 355 bp and an N50 of 506 bp.Comparisons with C.sinensis EST revealed not only the high-reliability but also the high-coverage of this transcriptome dataset.Sequence similarity analyses against seven public databases found 55,088 unigenes that could be annotated and assigned with gene ontology terms or putative metabolic pathways.Targeted searches using these annotations identified the majority of genes associated with several primary metabolic pathways and natural product pathways that are important to tea quality,such as flavonoid,theanine and caffeine biosynthesis pathways.Novel candidate genes of these secondary pathways were discovered.Thirteen unigenes related to theanine and flavonoid synthesis were validated.Their expression patterns in different organs of the tea plant were analyzed by quantitative real time PCR(qRT-PCR).In addition,12,242 SSRs distributed in unigenes were detected,including all repeat types from mononucleotide to hexanucleotide.Among them,the dinucleotide repeats are the main types,accounting for 63.78%of all the SSRs.The potential of the transcriptomic SSRs for further usage and research was assessed.2.Genome-wide SNPs discovered using RAD sequencing provide high-resolution species boundary and phylogenetic information for Camellia sinensis and its wild relativesUsing high-throughput genome-wide restriction site-associated DNA sequencing(RAD-Seq)technology,the simplified genome sequencing for 18 tea accessions including cultivated accessions from C.sinensis and wild accessions from four wild relatives/varieties were performed on Illumina platform.After data filtering,the effecctive tag sequences were 2.94±0.67 Gb on average.A total of 15,444 bi-allelic SNPs from 18 tea accessions were rapidly and cost-effectively generated after clustering of tag sequences and genotyping of nucleotide loci.Based on the identified genomic SNPs,all accessions were classified into six clusters corresponding to six Camellia species/varieties by phylogenetic,principle component and population structure analyses.It indicated the resultant genomis SNPs were suitable for high-resolution indetification of tested species/varieties and study of genetic relationship.Specifically,novel molecular evidence identified C.taliensis var.bangwei as a transitive tea plant possibly generated from interspecific hybridization of C.taliensis and C.sinensis var.assamica.Cultivated accessions exhibited greater heterozygosity than wild accessions,except for C.taliensis var.bangwei.A total of 1,521genic SNPs were identified from all 15,444 genomic SNP.Among them,1,058 unigenes were annotated with homologous Arabidopsis proteins,and 24 unigenes were identified to be related to secondary metabolic process.3.Whole genome sequencing and characterization of the draft genome of the tea plant (Camellia sinensis)and preliminary evolution analysisBefore the whole genome sequencing of tea plant,the genome survey on the cultivated tea clones of C.sinensis cv.Anhui1(AH1),C.sinensis cv.Tieguanyin(TGY)and C.sinensis cv.Shuchazao(SCZ)and one wild tea plant from C.taliensis(DXS)were carried out using Illumina sequencing technology before the sequencing of tea plant draft genome.There were105.3,110.8,205.1 and 159.8Gb of effective data from AH1,TGY,SCZ and DXS,respectively.Based on 17-mer analysis,the genome sizes of four tea plants ranged from 3to 3.3 Gb.The heterozygosities of them were between 1.5%and 2.2%,with the order of AH1>TGY>SCZ>DXS.The GC contents of four individuals were between 38.50%and39.92%.After preliminary assembly of four genomes,the results showed that the N50lengths of assembled contigs of them were all smaller than 700bp,and the N50 lengths of assembled scaffolds of them were all smaller than 3kb.Prediction of genomic SSRs(gSSR)from the preliminary assemblies of AH1 and TGY retrieved 563,680 and 545,520 results,respectively.Primer batch-designing successfully generated 262,807and 257,564 primer pairs for AH1 and TGY.C.sinensis cv.Shuchazao(SCZ)with the ralative lower heterozygosity in tested cultivated teas was selected for the material for whole genome sequencing of tea plant.The genome sequenging was performed using the whole genome shotgun strategy and high-throughput sequencing technology.The short-insert fragment libraries of 170-800 and the longt-insert fragment libraries of 2-40kb were construced and sequencing on Illumina Hiseq2000 platform for paired-end reads at an ultra depth.A total of 109 sequencing lanes were applied,which produced approximately 1,393Gb of high-quality clean data.In additon,the transcriptome sequencing of 8 tissues of SCZ were also performed,and 94.3Gb of clean data was obtained.After de novo assembly,the initiate contigs of 2.45Gb and the final 92,207 scaffolds of 2.98Gb was generated.The N50 length of assembled contigs and scaffolds were 33.4kb and 347.1kb,respectivly.Sequence comparison with C.sinensis EST from GenBank showed that the ESTs covered about 89.25%of the genomic region.The coverages of 3 randomly selected BACs individually determined through Sanger sequencing with the final assembly were 97.7%,100%and 94.8%,respectively.The tag sequences obtained from RAD-Seq of C.sinensis cv.Shuchazao were also used to assess the quality of assembled genomic regions flanking the restriction enzyme sites.It showed good agreement(96.79%)with the tea plant genome assembly.In the C.sinensis genome,approximately 56.06%of assembly was annotated to be repeats.TEs(transposable elements)accounted for 52.69%of the final assembly(94.0%of all repeats).Among them,the most aboundant TEs were indentified to be LTR(long-terminal repeat element).A total of 48,682 protein-coding genes were predicted,with an average gene length of 4,054bp bp and a mean of 3.3exons per gene.Based on GO,KEGG,SwissProt and InterPro databases,a total of 39,680 protein-coding genes were annotated.In addition,652 tRNA,3,450 rRNA,723 miRNA and 474 snRNA were also identified.In C.sinensis,a total of 33,922 genes were clustered into 17,701 gene families.Furthermore,3,415 unique gene families were identified to be tea-specific.There were7,171 gene families shared by tea plant,kiwifruit,Amborella and Eucalyptus.A phylogenetic tree was reconstructed from C.sinensis genome and other 11 sequenced genomes.It indicated that tea plant was grouped into Asterids,and had the closest relationship with Actinidiaceae.The time for the divergence of tea plant from kiwifriut was estimated at approximately 67.5 million years ago.Based on the tea plant draft genome,transcriptome and BAC library,the whole genomic DNA sequence of LAR that was one of key genes responsible for biosynthesis of catechins was identified and obtained,and the gene structure of LAR was also analyzed.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络