节点文献
基于Web挖掘的网页清洗技术
Web Page Cleaning Technology Based on Web Mining
【摘要】 随着互联网上信息的大量增多,Web挖掘技术越来越重要。而在Web挖掘过程中,基于Web的信息抽取的主要部分是如何去除网页中的噪音数据,它是Web数据的预处理的过程,这个预处理结果影响了Web挖掘的结果。在文中先分析了噪音数据的特点,然后根据实际观察提取规则并且用于模型统计的方法,去除噪音数据,抽取相关可利用的信息。
【Abstract】 With rapid expansion of information resources on the Internet increasingly,Web mining technology plays an important role.How to eliminate noisy information in web pages is a main part of information extraction based on Web mining.It is a preprocessing step in the Web mining.The result of Web mining lies on the step.In the paper,we firstly analyze the feature of noisy information.Then,based on our observation,using some extracting rules and statistic methods to eliminate noisy information and extract available information.
【基金】 国家自然科学基金资助项目(编号:90104021)
- 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2006年25期
- 【分类号】TP393.092
- 【被引频次】29
- 【下载频次】447