节点文献

基于文本内容的敏感词决策树信息过滤算法

Information Filtering Algorithm of Text Content-based Sensitive Words Decision Tree

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 邓一贵伍玉英

【Author】 DENG Yi-gui;WU Yu-ying;Information and Campus Network Management Center,Chongqing University;School of Computer Science,Chongqing University;

【机构】 重庆大学信息与网络管理中心重庆大学计算机学院

【摘要】 随着互联网的高速发展,各种各样的信息资源呈指数级增长,随之出现许多负面影响,需要构建一个安全健康的网络环境。为此,提出针对网页文本内容的敏感信息过滤算法(SWDT-IFA)。该算法不依赖词典与分词,通过构建敏感词决策树,将网页文本内容以数据流形式检索决策树,记录敏感词词频、区域信息以及敏感词级别,计算文本整体敏感度,过滤敏感文本。实验结果表明,SWDT-IFA算法具有较高的查准率和查全率,且执行时间能够满足当前网络环境的实时性要求。

【Abstract】 With the development of Internet,many negative effects come out as the exponential growth of various information resources,which means that a more secure and healthy network environment should be constructed right now.In order to solve this problem,this paper proposes a Sensitive Word Decision Tree for Information Filtering Algorithm(SWDT-IFA)for content-based Web pages. The algorithm takes no consideration of dictionary and word segmentation,builds the foundation on the sensitive words decision tree,lets the web text retrieval decision tree in form of data stream,records word frequency,regional information and sensitive level,and calculates the sensitive degree of the text to filter the sensitivity. Experimental results show that the SWDT-IFA algorithm has precision ratio and recall ratio,and low time complexity which can require the real-time demand of network environment.

  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2014年09期
  • 【分类号】TP391.1
  • 【被引频次】53
  • 【下载频次】973
节点文献中: 

本文链接的文献网络图示:

本文的引文网络