节点文献
杨属种质资源数据挖掘研究
Data Mining of Populus Germplasm Resources
【作者】 段旭良;
【导师】 冯秀兰;
【作者基本信息】 北京林业大学 , 森林经理学, 2008, 硕士
【摘要】 本文介绍种质资源的概念及其研究、保护的重要意义,在对数据挖掘以及知识发现的一般概念及方法进行分析的基础上,较为全面地总结了用于数据挖掘的聚类分析算法的类型及原理。在理论研究的基础上,通过对不同类型聚类算法的分析和比较,选取了k-means、k-mediods、FCM、CURE、SOM、GA等涉及划分方法、模糊方法、层次方法、人工智能和机器学习方法的6种主要聚类算法进行深入研究,并通过系统分析和设计、采用面向对象的程序设计方法应用C#语言编程实现,最终形成了一个较为通用的、集成多种方法的聚类分析软件系统。本文还对开发的软件进行了有效性测试,研究了聚类有效性评价函数。通过对二维随机数据点的聚类测试表明,程序能够满足一般情况下多维数据聚类要求,上述各聚类算法聚类是有效的,SD聚类有效性评价指标是科学有效的。研究发现:SOM自组织神经元网络相比其他算法效率更高、结果更稳定、效果良好,极具进一步深入研究的价值;SD有效性指标不仅可以用来评价聚类效果的好坏,而且可用于指导最佳聚类数的确定,具有进一步深入研究的价值。最后,本研究利用先期研究成果——基于聚类分析的数据挖掘软件系统——对杨属150个无性系的叶片因子数据进行了聚类分析,探索了杨属种质资源数据挖掘方法,初步完成了数据挖掘,抽取了部分典型的“类知识”,并讨论了“类知识”在遗传育种及良种繁育研究领域可能的应用,研究总体上取得了预期效果。
【Abstract】 This paper first discusses the significance of studying and protection on germplasm resources, and then expounds the basic theory and application of data mining studies in vast scopes.Based on the summarization, this paper selects 6 clustering methods such as k-means, k-mediods, FCM, CURE, SOM, GA and so on to lucubrate after comparing various clustering algorithms which are over partition, nesting, fuzzy, machine learning, artificial intelligence data mining methods. Meanwhile, this study implements the designed clustering software which is a multifunction-integrated and universal-used software system by using the object oriented programming language C#.This paper also researches into the cluster validity problem which is used to evaluate the validity of clustering. Two-dimension surface points’ clustering testing proves that the actualized software is availably, the SD validity index is logical and scientific. From this research and related studies we found that: i. Kohonen’s Self-Organizing Maps clustering method is more efficiently and more steadily, ii. The SD index is not only use for clustering-validity, but also for finding the optimal partitioning number of a data set. These are all worth to farther deeply study.At last, this study executes the data mining process, carries out the clustering-analysis on the leaf-factor data set of 150 Populus clones, and extracts several typical "group knowledge". In conclusion, this research acquires a series of significative result, reach it’s anticipate achievements.
【Key words】 Populus germplasm resources; data mining; clustering; cluster validity; C#;
- 【网络出版投稿人】 北京林业大学 【网络出版年期】2009年 01期
- 【分类号】S792.11
- 【被引频次】4
- 【下载频次】315