节点文献
基于堆栈稀疏自编码的K-均值聚类算法的种质评价
GERMPLASM EVALUATION BASED ON STACK SPARSE SELF-ENCODING K-MEANS CLUSTERING ALGORITHM
【摘要】 针对种质资源数据库构建过程中大量种质材料数据需要进行品质的分类的问题,提出堆栈稀疏自编码K-均值聚类算法对数据进行聚类,并将聚类结果利用已知品质标注的种质资源进行类别标注,从而达到对育种数据品质等级归类目的。区别于传统K-均值聚类算法,利用堆栈稀疏自编码网络进行关键数据特征提取,逐步降低样本维度,构建混合特征数据,将其作为K-均值聚类算法的初始中心,有效地避免了K-均值聚类算法中对初始中心选取敏感性问题。实验数据表明,该聚类算法准确率有明显提高。
【Abstract】 Aiming at the problem that a large amount of germplasm data needs to be classified in the process of constructing a database of germplasm resources,a stack sparse self-encoding K-means clustering algorithm was proposed to cluster the data. The clustering results were marked by the species quality resources with known quality,so as to achieve the purpose of classifying the quality data of the breeding data. Different from the traditional K-means clustering algorithm,the stack sparse self-encoding network was used to extract key data features. We gradually reduced the sample dimension and constructed mixed feature data as the initial center of the K-means clustering algorithm,effectively avoiding the sensitivity to the initial center selection in the K-means clustering algorithm. The experimental data showed that the accuracy of the clustering algorithm was significantly improved.
【Key words】 Clustering; Stack sparse self-encoding; Germplasm; Deep learning;
- 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2018年05期
- 【分类号】TP311.13
- 【下载频次】159