节点文献

基于图聚类的转录因子结合位点识别方法的研究

Study of Identifying Transcription Factor Binding Sites Based on Graphical Cluste

【作者】 王红岩

【导师】 马志强;

【作者基本信息】 东北师范大学 , 计算机应用技术, 2010, 硕士

【摘要】 转录因子结合位点是一段短的DNA序列,长度一般在10bp~30bp之间,它通常位于被调控基因上游的启动子区域中,转录调控蛋白与这些序列结合才能对基因的转录进行调控。因此识别转录因子结合位点成为构建转录调控网络的第一步。通常一种转录调控蛋白具有多个结合位点,这些位点在序列模式上具有相似性但又不完全相同,因此寻找一个转录调控蛋白的全部结合位点成为当今生物信息学领域最具挑战性的问题之一。本文提出一种利用图论知识对与一个转录调控蛋白结合的所有已知转录因子结合位点进行聚类从而识别未知转录因子结合位点的方法。通过对已知的转录因子结合位点进行聚类,我们可以把相似性最高的序列分到一组,对每一组中的序列构建位置特异性得分矩阵,从而得到多个位置特异性得分矩阵,把这些位置特异性得分矩阵组合在一起就形成所有已知的转录因子结合位点的混合位置特异性得分矩阵模型。我们用这个模型对训练集中的序列进行打分从而形成得分向量,用这些得分向量训练一个分类器,训练好的分类器就具有识别这种转录调控蛋白的结合位点的能力。理论上,我们的聚类方法不需要预先确定聚类数目而是根据序列之间的相似性自适应的调节聚类数,因此聚类效果比传统的聚类方法有较明显提高,另外,通过聚类得到的混合位置特异性得分矩阵模型的信息含量比单一的位置特异性得分矩阵高,因此用它给转录因子结合位点序列打分受随机事件的影响较小,分数更加可靠,训练分类器的效果更好,因此我们的方法比传统的方法更具优势。实验上,我们首先通过大肠杆菌转录因子结合位点测试我们的方法,结果识别效果比传统的位置特异性得分矩阵方法有明显提高。接着我们又对酵母的四个转录调控蛋白的结合位点序列进行试验,结果表明,使用我们的方法进行转录因子结合位点识别,在识别的敏感性和特异性上均比传统的位置特异性得分矩阵方法有较大提高,从而说明我们的方法在转录因子结合位点识别上是有效的。

【Abstract】 Transcription factor binding site (TFBS) is a short segment of DNA sequence. The length of TFBS is between 10 and 30 base pairs. It is usually located in the upstream of gene that is regulated by it. Transcription regulation proteins must bind to these sequences in order to regulate the transcription of the gene. So identifying the TFBS is the first step of constructing transcription regulation net. There are usually many binding sites to which one kind of transcription regulation protein binds. These sequences are usually similar but not same in the sequence mode. So identifying all binding sites of one transcription regulation protein becomes one of the most challenging problems today in the bioinformatics.The paper presents an approach that identify unknown binding site to which the transcription factor bind through a cluster method based on graph using all known binding sites to which the transcription factor bind. We can put the sequences that have the highest similarity into a group through clustering the known TFBS. Then we build a position-specific scoring matrix (PSSM) for each group of sequences so we will get some PSSMs. These PSSMs are put together to form a model called mix PSSM. We use this model to score each sequence in the training set and get a score vector of each sequence. Last we use these score vectors to train a classifier. The classifier that has been trained well has the ability to identify the binding site of the transcription factor.In the theory, our clustering method need not decide the count of cluster prior but modify adaptively the count according to the similarity of sequences. So the result of our clustering method increases more than the classic method. Otherwise, the information of mix PSSMs derived from cluster is higher than classic PSSM. So the score derived from mix PSSMs is disrupted less by random even. The score is more reliable and the result of training classifier is better. So our method has more advantages than classic method.In the process of experiment, first, we test our method by the TFBSs of Ecli-12. The results increase a lot relative to classic method of PSSM. Then, we do experiments for 4 kinds of TFBSs of yeast. The results indicate that the performance of identifying TFBS through our method increase more than classic method in sensitivity and specificity. So our method is effective in the identification of TFBS.

  • 【分类号】Q75
  • 【被引频次】3
  • 【下载频次】185
节点文献中: