节点文献

基于隐马尔可夫模型的二次k-均值基因序列聚类算法

A Double k-Mean Clustering Algorithm for Sequential Gene Data Based on the Hidden Markov Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吴君浩骆嘉伟王艳杨涛杨旭

【Author】 WU Jun-hao1,LUO Jia-wei1,WANG Yan1,YANG Tao1,YANG Xu2(1.School of Computer and Communications,Hunan University,Changsha 410082;2.School of Life Science,Hunan Normal University,Changsha 410081,China)

【机构】 湖南大学计算机与通信学院湖南师范大学生命科学学院 湖南长沙410082湖南长沙410082湖南长沙410081

【摘要】 本文提出了一种基于隐马尔可夫模型的二次k-均值聚类算法并实现了对基因序列数据的建模与聚类。算法首先引入了同源基因序列核苷酸比率趋向于一致的生物学特征来对基因序列数据进行初次k-均值聚类,然后利用第一次聚类结果训练出表征序列特征的隐马尔可夫模型,最后采用基于模型的k-均值方法再次聚类。实验结果表明,该算法是可行的,并且具有较好的聚类质量。

【Abstract】 A double k-mean clustering algorithm for modeling and clustering the gene sequence data is proposed by using the hidden Markov models(HMMs).First,the biological characteristics of four nucleotides ratio of homologous gene sequences is proposed to initial k-mean clustering on gene sequence data,and second,the first clustering results are utilized to train some HMMs which can denote sequence identities well.Finally,mode-based k-mean approach is adapted to clustering again.The experimental results show that the new algorithm is feasible and has comparatively better clustering quality.

【关键词】 隐马尔可夫模型基因序列建模k-均值聚类
【Key words】 HMMgene sequencesmodelingk-mean clustering
【基金】 湖南省自然科学基金资助项目(03jjy3095)
  • 【文献出处】 计算机工程与科学 ,Computer Engineering & Science , 编辑部邮箱 ,2007年03期
  • 【分类号】TP301.6
  • 【被引频次】2
  • 【下载频次】285
节点文献中: 

本文链接的文献网络图示:

本文的引文网络