节点文献

蛋白质-大分子相互作用及其分子疾病机制的生物信息学分析

Protein–macromolecule Interaction and Bioinformatics Analysis of the Molecular Disease Mechanism

【作者】 谢娟;

【导师】 刘士勇;

【作者基本信息】 华中科技大学 , Theoretical Physics, 2021, 博士

【摘要】 蛋白质通过和蛋白质、RNA、DNA相互作用在细胞生命过程中起关键作用,这些分子间结合的破坏可能扰乱细胞平衡从而致病。通过实验和计算方法识别并了解这些大分子间的结合有利于我们理解生命过程并为疾病的诊断和治疗提供思路。本文主要研究了蛋白质-大分子相互作用及分析了其分子疾病机制,蛋白质-大分子即蛋白质-蛋白质、蛋白质-RNA、蛋白质-DNA和蛋白质-RNA–DNA相互作用,并对蛋白质-RNA和蛋白质-DNA相互作用做了较为详细的研究,取得了初步进展。我们通过CLIP-seq和i RIP-seq高通量测序实验验证了RBPPred预测的CLIP1和DMD的确是RNA结合蛋白。我们对实验得到的结合位点数据进行分析,并开发了可以分析高通量测序数据的工具phd RBP。我们发现DMD及与之结合的RNA上的SNPs都可能与贝克尔肌营养不良、杜兴肌营养不良、扩张型心肌病3B和心血管表型有关。在13种癌症数据中,CLIP1和其他300个癌基因总是同时出现,且这300个基因中的123个与CLIP1相互作用。这些癌症可能与CLIP1及其与之相互作用的基因中的突变有关。尽管利用高通量测序可以获得特定蛋白质成百上千甚至上万的RNA结合位点,但是这个过程是费钱而且费时的,而且蛋白质-RNA的三维数据少,因此我们充分利用高通量测序产生的蛋白质-RNA结合位点数据及RNA二级结构,开发了基于模板的方法PRIME-3D2D来预测蛋白质-RNA相互作用的结合位点。在PDB和酵母基因组数据上测试表明PRIME-3D2D比其他结合位点预测软件表现更好。据研究报道,结合位点处的突变会影响蛋白质的稳定性从而影响蛋白质与其他分子的结合从而导致疾病。对于蛋白质-DNA,我们将PRIME2.0扩展到PDIME,预测了蛋白质-DNA的复合物结构,发现基于结构比对的方法比基于序列比对的方法可以找到更多的模板,DNA结构在蛋白质-DNA复合物结构的预测中起着重要作用。通过探索序列与结构的关系,我们发现许多具有不同序列的DNA具有相似的3D结构并执行相似的功能。这些数据及发现可以扩充分析复合物结构上的突变与疾病的关系,为将来构建有用的疾病分子模型打下基础。揭示不同基因导致相同疾病对理解和治疗疾病至关重要。为了探索蛋白质-大分子相互作用与分子疾病的关系,我们开发了3D2GENDIS将人类蛋白质-大分子复合物映射到基因组和疾病数据库,发现440807个突变发生在复合物的相互作用界面上,且这些突变中的大多数是致病的。我们将3D2GENDIS应用到COVID-19的复合物结构,发现相互作用界面上的大多数突变(86.9%)会降低蛋白质结构稳定性。三个由多个基因导致相同疾病的例子上发现突变可能会改变自由能进而改变蛋白质结构的稳定性。

【Abstract】 Proteins play a key role in the process of cell life by interacting with proteins,RNA,and DNA.Disrupting the binding between these molecules may disturb the balance of the cells and cause diseases.Identifying the interaction between these molecules through experimental techniques and computational methods will help us understand the life process and provide ideas for the diagnosis and treatment of diseases.In this article,we mainly study protein–macromolecule interactions and analyze their molecular disease mechanisms.Protein–macromolecules are protein–protein,protein–RNA,protein–DNA,and protein–RNA–DNA interactions.Among them,this article has done a more detailed study of protein–RNA and protein–DNA interactions,and has made preliminary progress.We identified CLIP1 and DMD are indeed RNA-binding proteins(RBPs)through CLIP–seq and i RIP–seq high–throughput sequencing experiments,CLIP1 and DMD are predicted as RBPs by RBPPred at first.We analyzed the binding site data obtained from the experiment and developed a tool phd RBP that can analyze high-throughput sequencing data.We found that the SNPs between DMD and its RNA partners may associate with Becker muscular dystrophy,Duchenne muscular dystrophy,Dilated cardiomyopathy 3B and Cardiovascular phenotype.Among the thirteen cancers data,CLIP1 and another 300 oncogenes always co-occur,and 123 of these 300 genes interact with CLIP1.These cancers may be related to the mutations in both CLIP1 and the genes it interacts with.Although the use of high-throughput sequencing can obtain thousands of RNA binding sites for a specific protein,this process is costly and time-consuming,and there are few protein–RNA three-dimensional data.Therefore,we make full use of the protein–RNA binding site data and RNA secondary structure generated by high-throughput sequencing,and developed a template-based method PRIME-3D2 D to predict the binding site of protein–RNA interaction.Testing on PDB and yeast transcription-wide data show that PRIME-3D2 D performs better than other binding sites predictor.Studies have reported that mutations at the binding site will affect the stability of the protein and thus affect the binding of the protein to other molecules,and then leading to diseases.For protein–DNA,we extended PRIME2.0 to PDIME to predict the structure of the protein–DNA complex.We found that structure-based methods can find more templates than sequence-based methods and DNA structure plays an important role in prediction of protein–DNA complex structure.By exploring the relationship of sequence and structure,we found that in protein–DNA interaction,numerous structures with dissimilar sequences have similar 3D structures and perform the same function.These data and findings can expand the analysis of the relationship between mutations in the complex structure and diseases,and lay the foundation for the construction of useful molecular models of diseases in the future.It is critical to understand and treat diseases by revealing that different genes cause the same diseases.In order to explore the relationship between protein–macromolecular interactions and molecular disease,we developed 3D2 GENDIS for mapping human protein–protein/RNA/DNA complex structures to the genome and disease databases.We found that 440807 mutations occur at the interface of the complex,and most of these mutations are pathogenic.We applied 3D2 GENDIS to analyze the complex structure of the COVID-19 and found that most of the mutations(86.9%)at the interaction interface will decrease the stability of the protein structure.By analyzing three examples which are composed of multi-genes but cause the same disease,we found that these mutations changed the free energy,which in turn changes the structure stability of the protein.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络