节点文献

基于MapReduce的新聚类算法在农业领域的应用——以柑橘红蜘蛛图像目标识别为例

Application of new clustering algorithm based on MapReduce in agriculture——A case study on image target recognition of Panonychus citri McGregor

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 卞云超司秀丽

【Author】 Bian Yunchao;Si Xiuli;Jilin Agricultural University;

【机构】 吉林农业大学

【摘要】 针对K-means聚类算法需要先给定k值,在一些应用场景中最优k值是未知的问题,提出基于评价机制的自适应K-means算法(SAK-means),并将该算法的核心步骤改写成Mapper/Reducer的形式,部署在Hadoop集群中。经过试验,该算法能够根据数据集的分布情况适当修正k值,特别适用于处理批量的、大尺寸的、最优k值非固定的聚类分析任务,并以批量的柑橘红蜘蛛图像目标识别为例进行验证,结果表明使用SAK-means算法无需给出最优的聚类中心数目,在一定范围内算法可以对聚类中心数目进行有效修正,对于实验中所选用的4幅图像,均可以达到100%的识别率与0%的误判率。进一步研究的方向是最优初始参数的选取,以及算法在集群中的扩展性与加速比。

【Abstract】 Aim to the problem that K-means clustering algorithm needs a preset kvalue and the best kvalue is unknown in some application scenarios,self-adaptive K-means algorithm(SAK-means)is proposed based on evaluation mechanism.The key step of SAK-means is rewritten to Mapper/Reducer,and SAK-means is also deployed in Hadoop cluster.Tested by experiments,SAK-means can appropriately modify k value according to distribution of the data sets,and it is especially applicable to process clustering analysis task with large bulk,large size and optimal unfixed kvalues,taking image target recognition of Panonychus citri McGregor as the case study,the results show that there is no need for SAK-means algorithm to provide the number of optimal cluster centers,because SAK-means effectively modify number of cluster centers within certain range,and its recognition rate can reach 100% while misjudgment rate is 0%.Further research will focus on selection of optimal initial parameters and expansibility and speed-up of algorithm in cluster.

【基金】 吉林省教育厅科学技术研究项目(201363);吉林省教育厅科学技术研究项目(201248)
  • 【文献出处】 中国农机化学报 ,Journal of Chinese Agricultural Mechanization , 编辑部邮箱 ,2016年09期
  • 【分类号】S126
  • 【被引频次】5
  • 【下载频次】126
节点文献中: