节点文献
基于无监督聚类的PU文本分类方法
PU-Oriented Text Classifier Based on Unsupervised Clustered Learning Algorithm
【作者】 张长利; 左万利; 彭涛; 赫枫龄; 彭钊; 邵慧勇;
【Author】 Zhang Changli~(1,2),Zuo Wanli~1,Peng Tao~1,He Fengling~1,Peng Zhao~3,and Shao Huiyong~3 1(College of Computer Science and Technology,Jilin University,Changchun 130012) 2(Shenyang College of Artillery Command,Shenyang 110162) 3(China National Petroleum Jilin Corporation,Songyuan 138000)
【机构】 吉林大学计算机科学与技术学院; 沈阳炮兵学院; 中国石油吉林油田分公司;
【摘要】 以正例(P)和未标识实例集(U)训练分类器的文本分类算法(PU文本分类)是解决某些机器学习中训练样本获取代价过大、尤其是反例样本较难获取的实际问题.而传统的分类算法大都需要正例和反例数据集才能取得良好的效果,因此要使用传统的分类方法来解决面向PU的分类问题,U集中可信反例的提取是分类器能够取得良好效果的关键.提出了有效的可信反例提取算法(基于聚类的可信反例提取算法)——CBRN,并对已有的PU文本分类算法进行了改进,并提出了SPY-SVM算法.实验表明,该方法比目前其他的面向PU的文本分类方法具有更高的准确率和召回率.
【Abstract】 Presented here is a new PU-oriented text classifier learning algorithm.It solves problems in machine learning when no labeled negative documents are available in the training example set or when negative examples are very difficult to collect.Traditional classification algorithm can’t obtain good performance without numerous positive and negative training dataset.When using traditional classifier to solve PU-oriented text classification,the key is the extraction of reliable negative example with the CBRN(cluster based reliable negative example extraction) algorithm.Existing classification is then improved,which builds a set of classifiers by iteratively applying the SPY-SVM algorithm. Experimental results on the Reuter data set show that this method outperforms other algorithms in terms of Fl-measure.
- 【会议录名称】 第二十五届中国数据库学术会议论文集(二)
- 【会议名称】第二十五届中国数据库学术会议
- 【会议时间】2008-10-24
- 【会议地点】中国广西桂林
- 【分类号】TP18
- 【主办单位】中国计算机学会数据库专业委员会