节点文献

结合基因组和转录组学解析斑马演化模式

Combining Genomics and Transcriptomics to Investigate Zebra Evolution Model

【作者】 萨如拉

【导师】 芒来;

【作者基本信息】 内蒙古农业大学 , 动物遗传育种与繁殖, 2018, 博士

【摘要】 斑马(Zebra)是马属(Equus)动物成员之一,主要分布在非洲草原,包括平原斑马、细纹斑马和山斑马三个亚种,其中平原斑马数量较多,其它两个亚种已被列入到濒危物种名单。马属动物基因组及进化研究较多,但对斑马基因组方面的研究屈指可数,只有线粒体基因组及少数基因组重测序数据。目前,大规模测序技术由于其成本低,通量高等优势已被广泛应用于多个哺乳动物基因组的测定及分析。马属中马和驴基因组已被测序,并拼装程度已达到染色体及亚染色体水平,这使马和驴基因组间的比较和溯源研究成为可能。平原斑马虽有基因组序列相关研究,但这些读段水平的序列不足以分析全基因组范围内的遗传变异及系统进化研究。因此,为马属动物的系统发育及全面的进化分析成为可能,我们对平原斑马基因组进行从头测序。另外,在现有生物学数据库中斑马转录水平的研究基本空白,因此,我们对平原斑马不同组织进行转录组测序,进而通过同源基因的进化速率,探索斑马适应性进化的遗传机制。通过基因组和转录组测序的两部分研究内容,得出的主要结论如下所述:(1)利用Illumina Hiseq/Miseq平台对平原斑马基因组进行测序,获得570733246条质控后的序列。将质控后的序列与马参考序列(EquCab2.0)进行比对分析,获得26032374个SNPs和687552个InDel,主要分布在基因间区。经Newbler拼接后,获得总序列长度为2.36Gb的基因组,其Contig和Scaffold N50分别为43.7Kb和1.45Mb。斑马基因组中注释的重复序列含量为42.61%,蛋白编码基因数为22732。(2)通过种间同源基因家族推断马和非马类的分化时间为9.2-25.5百万年前,斑马和驴的分化时间为6.7-22.6百万年前。同时发现斑马基因组上有872个扩张的基因家族和1750个快速进化的基因,这些基因主要富集在锌指、转录及转录调控、鞘脂代谢、蛋白复合支架、序列特异性DNA结合、细胞外刺激的反应、T辅助2型免疫应答调节、由RNA聚合酶II启动子的转录调控和转录因子活性等生物学功能。(3)利用数据库中可用的马科动物基因组序列,并检测SNP后构建马科动物亲缘关系树发现,平原斑马与quagga和bohmi聚在一起,并与细纹斑马和山斑马构成单个斑马进化支。PSMC推断的马科动物种群历史表明,斑马种群规模轨迹与马和驴有所不同,这提示着美洲、欧亚和非洲具有非同步的复杂生态动态变化。(4)将从头拼接的平原斑马基因组序列草图与马染色体序列进行比对,进行共线性分析的同时共检测到2207个重排事件,包括1664未知插入、2重复插入、204倒位和337易位。并且这些重排区域富含LINE/L1、Satellite、LTR/ERV1和SINE/tRNA等重复类型。另外,通过同源序列比对,确定平原斑马2号染色体的着丝粒特征序列,即SAT2pl和SATEC卫星序列,与马的着丝粒特征序列一致。(5)利用Illumina Hiseq x Ten平台对平原斑马5个组织进行转录组测序,最终各组织获得6779534289462708条高质量序列。用Trinity从头拼接获得752562个转录本和482219个unigenes,其平均长度分别是1127bp和631bp,N50分别是2719bp和954bp。经注释,至少在一个数据库中获得注释信息的unigenes数量为243758条(50.54%)。从平原斑马unigenes中共搜索到69096个SSRs。(6)我们以FPKM>0.3为基因的表达阈值,在5个组织中筛选出的基因表达数目在78181216298之间,其中共有表达的有23672个,并且骨骼肌和心脏的共有表达基因表达水平最相似。特有表达基因在肺脏和肾脏中居多。(7)通过差异基因分析,在肾脏和骨骼肌间差异表达的基因最多,其次是肾脏和骨骼肌。经注释,TNNT2和TNNI3等基因在心脏中显著高表达;ALB、CYP2D和UGT等基因在肝脏中显著高表达;TNNC2、TNNI2和TNNI1等基因在骨骼肌中显著高表达;UMOD基因在肾脏中显著高表达;HSPA18和SFTPC等基因在斑马肺脏中显著高表达。(8)基于直系同源基因的进化速率,在斑马基因组上检测到877个受正选择的基因,其中284个受显著正选择(P<0.05),270个是极显著正选择的基因(P<0.01)。通过GO和KEGG富集分析发现,这些正选择的基因主要参与免疫、神经、血管生成、紫外线保护和胰岛素分泌等有助于适应热带气候的代谢通路和生物学功能相关分类。(9)通过对斑马不同组织的小RNA进行测序,获得1406188916662677条高质量序列数据,主要分布在21-23nt长度范围内。将其与miRBase中马的已知序列比对,获得204个保守miRNA和274个miRNA前体,同时新预测出78个成熟miRNA和83个miRNA前体。已知和新miRNA的首位碱基具有U碱基偏好性。(10)以TPM≥0.1为miRNA表达阈值,发现在斑马组织中中丰度和高丰度表达的miRNA占用比例较高。三个组织间差异表达的miRNA有127个,其中心脏和肝脏差异的有85个,心脏和骨骼肌差异的有25个,肝脏和骨骼肌间差异的有86个。对282个miRNA预测出34205个靶基因,并差异表达miRNA的靶基因主要富集在分子功能、蛋白结合、细胞组分、细胞过程、代谢过程等GO功能分类及Ras信号通路、JAK-STAT信号通路、神经营养因子的信号转导通路及胰岛素抵抗等KEGG代谢通路。本研究结果对马属动物分子生物学研究提供序列数据资源,并为日后的深入研究奠定基础。

【Abstract】 Zebra is one of the members of Equus,which mainly distributed in the African grassland and including three subspecies of plains zebra,grevyi and mountain zebra.Among them,there are a large number of plain zebra,and the other two subspecies have been listed on the endangered species list.There are many studies on Equus genome and evolution,but only a few studies on the zebra genome,including mitochondrial genome and a few genome re-sequencing data.At present,large-scale sequencing technology has been widely used in the sequencing and analysis of multiple mammalian genomes due to its low cost and high throughput.The genome of horse and donkey in equus was sequenced and the assembly has reached the chromosomal and sub-chromosomal level,which makes the comparison and traceability between horse and donkey genome possible.Although plains zebra has genomic sequence related studies,these read-level sequences are not enough to analyze genome-wide genetic variation and phylogenetic studies.Therefore,in order to make phylogenetic analysis and comprehensive evolutionary study of equus possible,we performed de novo sequencing on the plain zebra genome.In addition,the research on zebra transcriptional level in existing biological databases is basically blank.Therefore,we performed transcriptome sequencing on different tissues of the plains zebra,and then explored the genetic mechanism of zebra adaptive evolution through the evolutionary rate of orthologous genes.Through the two parts of the genome and transcriptome sequencing,the main conclusions are as follows:(1)We obtained 570733246 clean reads for zebra genome using Illumina Hiseq and Miseq platform.Then identified 26032374 SNPs and 687552 InDels in zebra genome comparing with horse(Equ2.0),and they mainly distributed in the intergenic region.After assembled with Newbler,we obtained 2.36 Gb of total sequence length,and its Contig and Scaffold N50 were 43.7Kb and 1.45 Mb,respectively.The repeat content of zebra genome is 42.61% and the number of protein coding genes is 22732.(2)We inferred the differentiation time for caballine and non-caballine lineage was 9.225.5 Mya,and 6.722.6 My for zebra and donkey using homologous gene tree.The zebra genome has 872 expanded gene families and 1750 rapidly evolving genes,which were mainly enriched in zinc finger,transcription,regulation of transcription,sphingoid metabolic process,protein complex scaffold,sequence-specific DNA binding,cellular response to extracellular stimulus,regulation of T-helper 2 type immune response,regulation of transcription from RNA polymerase II promoter and transcription factor activity.(3)The affinity relationship tree using SNP from Equidae that available in public database has revealed that plains zebra clustered with quagga and bohmi,and formed zebra monophyletic branch together with Greyvi and Mountain zebra.The PSMC curve showed different population trajectory for zebra from that of horse and donkey,suggesting asynchronous and complex ecological dynamic changes in America,Eurasia and Africa.(4)Through comparing plains zebra draft genome with horse genome,2207 rearrangement events were detected,including 1664 insertion of unknown origin,2 inserted duplication,204 inversion and 337 relocation.And these rearrangement regions are rich in LINE/L1,Satellite,LTR/ERV1 and SINE/tRNA repeat types.We also detected SAT2 pl and SATEC satellites were centromere sequence of plains zebra chromosome 2 which is consistent with the centromere sequence of the horse(5)We obtained 6779534289462708 high quality reads through RNA-seq with Illumina Hiseq x Ten for 5 different tissues of zebra.After assembling with Trinity,we obtained 752562 transcripts and 482219 unigenes.The average length of them was 1127 bp and 631 bp,and the N50 was 2719 bp and 954 bp,respectively.After annotation,the number of unigenes annotated in at least one database is 243758(50.54%).We detected 69096 SSRs on plains zebra unigenes.(6)We used FPKM>0.3 as the gene expression threshold,has screened 78181216298 expressed genes for 5 tissues,containing 23672 common expressed genes with most similar expression level in skeletal muscle and heart.The specific expression genes were most in the lung and kidney.(7)Through DEG analysis,the most differentially expressed genes were found in kidney and skeletal muscle,followed by kidney and skeletal muscle.After annotation,TNNT2 and TNNI3 were highly expressed in the heart;ALB,CYP2 D and UGT were highly expressed in the liver;TNNC2,TNNI2 and TNNI1 were highly expressed in the skeletal muscles;UMOD was highly expressed in the kidneys;HSPA18 and SFTPC were highly expressed in the lung.(8)Based on the substitution rate of orthologous genes,we detected 877 positively selected genes in zebra genome,of which 284 were significantly positively selected(P<0.05)and 270 were extremely significant positively selected genes(P<0.01).Through GO and KEGG enrichment analysis,these positively selected genes were mainly involved in immunity,nerve,angiogenesis,ultraviolet protection and insulin secretion,which could contribute to adaptation in tropical climate.(9)Through small RNA sequencing on different tissues of zebra,we obtained 1406188916662677 high quality reads mainly distributed in 21-23 nt length range.Comparing with horse known sequences in miRBase,204 conserved miRNA and 274 miRNA precursors were obtained.Meanwhile,78 mature miRNA and 83 miRNA precursors were predicted.Known and novel miRNA all have U base preference at its first base.(10)Used TPM ≥ 0.1as miRNA expression threshold,moderate and high abundance miRNA in zebra tissues occupy a higher proportion.There were 127 differentially expressed mi RNAs between three tissues of zebra,including 85 between heart and liver,25 between heart and skeletal muscle,and 86 between liver and skeletal muscles.34205 target genes were predicted for 282 miRNAs,and the target genes of differentially expressed miRNAs were mainly concentrated in GO function classification,including molecular function,protein binding,cell component,cell process and metabolic process.And KEGG pathways,including Ras signaling pathway、Jak-STAT signaling pathway、Neurotrophin signaling pathway and Insulin resistance.The results of this study provide sequence data resources for studies on molecular biology of equus and lay the foundation for further research in the future.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络