节点文献

一种基于自适应Markov模型的中文垃圾邮件过滤方法

Method for Filtering Chinese Spam Based on the Adaptive Markov Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李劲岳昆杭菲璐

【Author】 LI Jin~1 YUE Kun~2 HANG Fei-lu~1 (School of Software,Yunnan University,Kunming 650091,China)1 (School of Information Science and Engineering,Yunnan University,Kunming 650091,China)2

【机构】 云南大学软件学院云南大学信息学院

【摘要】 现有的中文垃圾邮件过滤方法将邮件模型化为词包,针对词包模型存在中文分词困难、忽略语境中上下文的相关性等不足之处,将邮件看作字符序列、无需对其进行分词,建立描述不同类别邮件的序列统计特征的自适应Markov模型,进而构造出基于字符序列的贝叶斯邮件过滤器。以CCERT垃圾邮件作为语料集,对提出的自适应Markov模型过滤方法进行了测试,实验结果表明本文的方法具有较高的查全率和查准率,性能优于现有的方法。

【Abstract】 Existing methods for filtering Chinese spasm transfer mails to word packages.By these methods,it is hard to make stemming and obtain the associations implied in linguistic contexts.Therefore,we look upon mails as sequences of characters,instead of making stemming on the mails.We establish the adaptive Markov model for describing sequential statistic characteristics of various types of mails,and then develop the Bayesian mail filter based on character sequences.Adopting CCERT spam as the data set,we test the method proposed in this paper.Experimental results show the high recall and precision of our method and the high performance compared to existing methods.

【基金】 云南大学理(工)科青年科研基金(No.2005Q023C);云南省自然科学基金项目(No.2005F0009Q);国家自然科学基金项目(No.60763007)
  • 【会议录名称】 第二十五届中国数据库学术会议论文集(一)
  • 【会议名称】第二十五届中国数据库学术会议
  • 【会议时间】2008-10-24
  • 【会议地点】中国广西桂林
  • 【分类号】TP393.098
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络