节点文献
基于Tabu的Deep Web特征选择算法
Feature selection of deep web based on Tabu
【摘要】 Deep Web分类的小样本、高维特征的特点限制了分类算法的选择,影响分类器的设计和准确度,降低了分类器的"泛化"能力,出现分类器"过拟合",所以需要进行特征选择,降低特征的维数,避免"维数灾难"。目前,没有Deep Web特征选择自动算法的相关研究。通过对Deep Web分类的特征选择进行研究,提出了基于类别可分性判据和Tabu搜索的特征选择算法,可以在2的时间复杂度内得到次优的特征子集,减小了分类器设计的难度,提高了分类器分类准确率。根据特征选择前后的特征集,利用KNN分类算法进行Deep Web分类,结果表明提高了分类器的分类准确率,降低了分类算法的时间复杂度。
【Abstract】 Classification of deep web has characteristic of small sample and high dimensional,which restricts choice of classification algorithm and makes the classifier hardly design,also lower the" Generalization Ability "and makes the classifier" overfitting." Feature selection is necessary to avoid"curse of dimensionality." There is no research about automatic classification algorithm at present.Feature selection algorithm of deep web based on Tabu search algorithm and separative criterion is put forward through the research about feature selection,which can quickly find feature subset in the time complexity of 2.Feature selection algorithm based on Tabu and separative criterion makes design of classifier easily,also increases accuracy of classifier.Experimentation based on classifier indicates that feature selection algorithm based on Tabu and separative criterion increases accuracy of classifier and reduces computation complexity.
【Key words】 feature selection; tabu search algorithm; deep web; information retrieval; classification algorithm; classifier;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2008年13期
- 【分类号】TP391.1
- 【被引频次】1
- 【下载频次】112