节点文献
大型数据库中数据挖掘算法SLIQ的研究及仿真
Simulation Study on Data Mining Algorithm SLIQ in Large Databases
【Author】 GUO Xin-yu, LIANG Xun(Institute of Computer Science and Technology,Peking Univ ersity,Beijing 100871 ,China)
【机构】 北京大学计算机科学技术研究所;
【摘要】 大型数据库中的数据挖掘是目前数据挖掘领域重要的前沿课题。SLIQ算法是处理大型数据库数据的有效算法之一。该算法分为三个阶段:在预处理阶段,对数值属性进行预排序;在树的构建阶段采用了树的宽度优先增长策略,使得SLIQ能够处理海量的磁盘数据;在树的修剪阶段,采用了一种基于最小描述长度原理的修剪算法。实验结果表明,SLIQ能够较快且准确地处理海量的数据。
【Abstract】 Data mining in large databases is an important advanced subject in the field of data mining at present. SLIQ is one of the effective algorithms that can handle data sets in large databases. The algorithm is composed of three phases: in the pre-processing phase, a pre-sorting technique is applied on the numerical attributes; In the tree-building phase, a bread-first tree growing strategy is used to enable SLIQ to handle large disc-resident data sets; A pruning algorithm based on the principle of Minimum Description Length is applied in the tree-pruning phase. The simulation results of experiments show that SLIQ can handle large data sets fast and correctly.
【Key words】 large database; data mining; SLIQ algorithm; gini index;
- 【会议录名称】 2004年中国管理科学学术会议论文集
- 【会议名称】2004年中国管理科学学术会议
- 【会议时间】2004
- 【分类号】TP311.13
- 【主办单位】中国优选法统筹法与经济数学研究会