节点文献

基于多示例多标签支持向量机不平衡网页分类

Imbalanced Web-page Classification Based on Multi-instance Multi-label Support Vector Machine

【作者】 唐磊

【导师】 李村合; 庄涛;

【作者基本信息】 中国石油大学(华东) , 计算机技术(专业学位), 2017, 硕士

【摘要】 随着Internet的普及,网络已经成为人们获取信息的主要途径,为了方便人们从海量网页中获取有用的信息,一种网页自动分类技术应运而生。鉴于多示例多标签(MIML)框架在歧义性学习方面独特的先天优势,以及支持向量机(SVM)卓越的学习能力,二者融合算法目前已成为机器学习领域的研究热点,但是这两者在处理不平衡网页的时候会有所欠缺。介绍了网页分类过程及其相关技术,描述了MIML框架理论,阐述了SVM发展历程、理论原理,并着重分析了在MIML框架下发展而来的MIMLSVM和MIMLSVM+算法。在样本集中,经常会出现某一类样本明显多于另一类样本的情况,这就得使样本集不够均衡。针对MIMLSVM算法在这种不平衡样本下分类效果差的问题,先使用过采样方法对样本集进行预处理,让样本集变得更加均衡,最终提高了分类准确率。现实生活中,无标签的的样本有很多,有标签的样本却非常少,数量众多的无标签样本可以提供整个样本空间的分布状况,进而弥补了少量有标签样本缺点。针对MIMLSVM+算法在这种不平衡样本下分类效果差的问题,提出了利用渐进直推向量机思想来处理无标签的样本,进而提高了分类准确率。最后,将改进后的训练算法应用到网页分类系统中,并对改进算法进行了实验对比和性能分析。实验数据表明,本文算法具有更高的分类效率和准确率。

【Abstract】 With the popularity of Internet,the network has become the main way for people to obtain information.In order to help people get useful information from massive web pages,web page automatic classification technology emerges as the times require.In view of multi instance multi label(MIML)in the framework of the ambiguity of learning a unique advantage,and the support vector machine(SVM)excellent learning ability,the two fusion algorithm has become a research hotspot in the field of machine learning.However,the combination of the two is not enough to deal with the imbalance of the web page.Introduces the basic process of web page classification and its related technologies,describes the framework of MIML theory and algorithm,and discusses the principle of SVM development history,theory,and analyzes the development of MIMLSVM and MIMLSVM+ algorithm under the framework of MIML.Because the sample will be a sample of significantly more than other types of sample collection,according to the MIMLSVM in this sample are unbalanced classification problems,proposed to preprocess the sample by random sampling method improved,reducing the impact of unbalanced sample the network model.In real life,no labeled samples,but there are very few labeled samples,the distribution of a large number of unlabeled samples can provide the whole sample space,and then make up a small amount of labeled samples is difficult to describe the space distribution of samples,shortcomings,and can improve the performance of the classifier.In order to solve the problem of poor classification performance of MIMLSVM+ in this kind of unbalanced samples,a direct vector machine based on the two programming is proposed to deal with unlabeled samples,and the classification accuracy is improved.Finally,the improved training algorithm is applied to the web page classification system.The experimental results show that the proposed algorithm has higher classification efficiency and accuracy.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络