节点文献
基于语义的文本数据流概念漂移检测算法
Semantic-based Concept Drift Detection Algorithm for Text Data Stream
【摘要】 文本数据流中概念的频繁漂移导致有效信息不足,从而使得漂移检测和数据流分类准确率下降。针对该问题,引入潜在狄利克雷分布模型并考虑文本数据流隐含的语义信息,提出一种新的概念漂移检测算法。计算相邻模块中词和主题特征空间的语义相似度,其中主题的相似度根据主题-单词概率分布进行评估,当2个特征空间相似度都较低时判断为发生概念漂移。实验结果表明,与DDM、CDRDT、DWCDS、HDDM-W-Test和REDLLA算法相比,该算法对文本数据流中概念漂移的检测性能均有所提升,尤其在概念频繁漂移时可以显著减少漏检数量。
【Abstract】 In text data stream,frequent concept drifts result in the poor effective information,thus the accuracy rates of drift detection and stream classification are lower. To address this problem,by introducing Latent Dirichlet Allocation(LDA) model and considering the semantic information of text data stream,this paper proposes a newconcept drift detection algorithm. It calculates the semantic similarities of both word and topic feature spaces between adjacent modules,in which the similarity of topics is evaluated by the probability distribution of topic-word. It is considered that concept drifts occur when the similarities are lower in these two spaces. Experimental results showthat,compared with DDM,CDRDT,DWCDS,HDDM-W-Test and REDLLA algorithms,the proposed algorithm can improve the performance in the concept drift detection. Especially,it can significantly reduce the missing drifts when concept frequently drifts.
【Key words】 concept drift; semantic; drift detection; Latent Dirichlet Allocation(LDA) model; text data stream classification;
- 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2018年02期
- 【分类号】TP391.1
- 【被引频次】6
- 【下载频次】189