节点文献

人类基因组非冗余Exon/Intron数据库的构建

CONSTRUCTION OF HUMAN NON-REDUNDANT EXON/INTRON DATABASE

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 罗冬梅金鹰邓小元刘海

【Author】 LUO DongMei1,JIN Ying1,DENG XiaoYuan1,LIU Hai2(1.College of Biophotonics,South China Normal University,Guangzhou 510631,China;2.School of Computer,South China Normal University,Guangzhou 510631,China)

【机构】 华南师范大学生物光子学研究院华南师范大学计算机学院

【摘要】 以Homo.sapiensRefSeq作为原始数据库来构建EID(Exon/Intron Database)可以克服GenBank所带来的冗余问题.通过分析RefSeq基因组数据库中每个CDS(Coding Sequence,编码序列),获得构建EID的相关的数据(基因的定义、基因标识符、基因序列、蛋白质标识符、蛋白质序列、外显子和内含子的数量、大小、总数、非翻译区(UTR)内含子、内含子相位、内含子剪切位点模式).结果表明,人类24条染色体(22条常染色体和2条性染色体,共计2 870 827355 bps)中含有32 157个基因标识符(gene blocks),其中7 398个基因为假基因,4 014个基因发生了可变剪切(Al-ternative Splicing,AS),15 533个基因含有CDS内含子,765个基因含有UTR内含子,2 585个基因不含有内含子,其他的为异常基因.

【Abstract】 The exon/intron database(EID) is redundant when it is constructed based on GenBank records.In order to overcome this shortcoming,a non-redundant EID is derived from Homo.sapiens RefSeq(Reference Sequence) database.After analysing each CDS(Coding Sequence) field in original RefSeq database,the data related to eukaryotic genes(definition line,gene_id,gene sequence,protein_id,protein sequence,number of exon(s) and intron(s),size of exon and intron,sum of exons and introns,intron in UTR,phase of intron,and pattern of splice site) are collected into EID.All of human chromosomal sequences(total 2 870 827 355 bps) are parsed and we obtain 32 157 gene blocks.In there genes,there are 7 398 pseudo genes,4 014 alternative splicing genes,15 533 genes with intron in CDS,765 genes with intron in UTR,2 585 genes without intron,and other imperfect genes.

【基金】 国家自然科学基金专项项目/科学部主任基金项目(30940020);国家自然科学基金项目(30470495)
  • 【文献出处】 华南师范大学学报(自然科学版) ,Journal of South China Normal University(Natural Science Edition) , 编辑部邮箱 ,2010年04期
  • 【分类号】Q75
  • 【被引频次】1
  • 【下载频次】143
节点文献中: 

本文链接的文献网络图示:

本文的引文网络