节点文献

基于最大熵选取示例的增量决策树归纳

Sample Selection Based on Maximum Entropy for Incremental Induction of Decision Trees

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 闫建辉王熙照隋春荣王硕苑俊英

【Author】 YAN Jian-hui1,WANG Xi-zhao1,SUI Chun-rong2,WANG Shuo1,YUAN Jun-ying1(1.College of Mathematics & Computer Science,Hebei University,Baoding,Hebei 071002,China;2.Xingtai University,Xingtai,Hebei 054001,China)

【机构】 河北大学数学与计算机学院邢台学院河北大学数学与计算机学院 河北保定071002河北保定071002河北邢台054001

【摘要】 设A是一训练集,B是A的一个子集,B是选择A中部分有代表性的示例而生成的。得到了这样一个结论,即对于适当选取的B,由B训练出的决策树其泛化精度优于由A训练出的决策树的泛化精度。进一步,设计实现了一种如何从A中挑选有代表性的示例来生成B的算法,并从数据分布和信息熵理论角度分析了该算法的设计原理。

【Abstract】 Suppose that A is a training set and B is a subset of A.B is generated by selecting some representative samples from A.This paper draws such a conclusion that,for appropriately selected B,the generalization capability of decision tree trained on B is better than the decision tree trained on A.Furthermore,an algorithm of generating B by selecting representative samples from A is designed.And from the viewpoints of data distribution and information entropy,the algorithm is analyzed.

【基金】 国家自然科学基金资助项目(60473045,60573069)。
  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2006年35期
  • 【分类号】TP18
  • 【被引频次】7
  • 【下载频次】174
节点文献中: 

本文链接的文献网络图示:

本文的引文网络