节点文献
基于最大熵选取示例的增量决策树归纳
Sample Selection Based on Maximum Entropy for Incremental Induction of Decision Trees
【摘要】 设A是一训练集,B是A的一个子集,B是选择A中部分有代表性的示例而生成的。得到了这样一个结论,即对于适当选取的B,由B训练出的决策树其泛化精度优于由A训练出的决策树的泛化精度。进一步,设计实现了一种如何从A中挑选有代表性的示例来生成B的算法,并从数据分布和信息熵理论角度分析了该算法的设计原理。
【Abstract】 Suppose that A is a training set and B is a subset of A.B is generated by selecting some representative samples from A.This paper draws such a conclusion that,for appropriately selected B,the generalization capability of decision tree trained on B is better than the decision tree trained on A.Furthermore,an algorithm of generating B by selecting representative samples from A is designed.And from the viewpoints of data distribution and information entropy,the algorithm is analyzed.
【关键词】 样例挑选;
信息熵;
模糊决策树归纳;
泛化精度;
【Key words】 sample selection; information entropy; fuzzy decision tree induction; generalization capability;
【Key words】 sample selection; information entropy; fuzzy decision tree induction; generalization capability;
【基金】 国家自然科学基金资助项目(60473045,60573069)。
- 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2006年35期
- 【分类号】TP18
- 【被引频次】7
- 【下载频次】174