节点文献
基于卡方方法及对称不确定性的网络流量特征选择方法
Network Traffic Feature Selection Method Based on Chi-square Method and Symmetric Uncertainty
【摘要】 对网络流量数据进行分类时,由于网络流量具有多个类别,并且各类样本数量不均衡,故在利用机器学习进行分类时,会导致分类的模型的性能降低,致使样本被误分为样本数量多的类别,进而致使样本数量较少的类别(小类别)的召回率过低。针对该问题,提出一种基于卡方方法及对称不确定性网络流量特征选择方法。该方法首先计算特征与类之间的加权卡方值,选择卡方值较大的特征组成候选特征子集,然后根据特征与所有类之间的对称不确定性进一步筛选特征集。在Moore网络流量数据集上进行实验,得到的实验结果证明,通过该方法选择的特征对网络流量数据进行分类,在保证准确率高的前提下也得到了较高的小类召回率,减轻了数据不均衡问题带来的不良影响。
【Abstract】 When classifying network traffic data,because network traffic has many categories and the number of samples is not balanced,the performance of classification model will be reduced while machine learning is used to classify network traffic data. As a result,samples are mistakenly classified into categories with a large number of samples,and the recall rate of smaller categories(small categories) is too low. In order to solve this problem,a chi-square method and a symmetric uncertain network traffic feature selection method are proposed. Firstly,the weighted chi-square values between the features and the classes are calculated in the method;and the features with larger chi-square values are selected to form a candidate feature subset. Then the feature sets are selected according to the symmetry uncertainty between the features and all classes. The experimental results on the Moore network traffic data set show that the classification of the network traffic data by the selected features of the method can also obtain a higher recall rate of small classes on the premise of high accuracy. The negative impact of the data imbalance is mitigated.
【Key words】 data imbalance; network traffic; relative uncertainty; recall rate;
- 【文献出处】 长春理工大学学报(自然科学版) ,Journal of Changchun University of Science and Technology(Natural Science Edition) , 编辑部邮箱 ,2019年02期
- 【分类号】TP393.06
- 【被引频次】3
- 【下载频次】146