节点文献

基于向量空间模型的基因序列聚类及仿真实验

Cluster analysis and emulation of gene sequences based vector space model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张东生季超

【Author】 (Computing Center, Henan University, Kaifeng 475004, China) ZHANG Dong-sheng JI Chao

【机构】 河南大学计算中心

【摘要】 聚类算法广泛应用于生物信息学数据分析中,是基因序列和表达数据分析研究的主要技术之一。提出了一种基于向量空间模型的基因序列聚类分析算法。首先利用DNA序列的结构特征,将多个DNA序列构成序列集。结合向量空间模型算法,计算DNA序列集中两两序列之间的相似度矩阵,并选取适当的阈值对相似度矩阵作截集处理,从而得到最终的聚类结果。基于DNA序列数据的仿真实验结果表明,该算法在基因序列的分析中是实用、有效的,并且具有算法简明、语义准确、向量维数可控等优点。

【Abstract】 Clustering algorithms, which is one of the main techniques for analyzing gene sequences and expression data, are widely applied in the research of bioinformatics data. A clustering algorithm for gene sequences analysis based on vector space model is proposed in this paper. Firstly, according to the structure characteristics of gene sequences, different bases in DNA are used to construct the DNA sequences which consist of the DNA sequences set. Then the similarity matrix between DNA sequences is computed using the vector space model algorithm. The final cluster results are obtained by choosing the proper threshold for the similarity matrix over cuts. Simulation results on the DNA sequences data have shown the vector space model algorithm is very feasible, efficient in gene sequences analysis. The presented algorithm has the advantages of conciseness, semantic accuracy and the controllable dimension of the vector.

  • 【文献出处】 微计算机信息 ,Microcomputer Information , 编辑部邮箱 ,2010年16期
  • 【分类号】TP18
  • 【下载频次】98
节点文献中: 

本文链接的文献网络图示:

本文的引文网络