节点文献
Web主题关联知识自学习算法
Discovery of Web Topic-Specific Association Rules
【摘要】 <正> 1 概述面向主题的Web信息搜索和挖掘是当前的一个研究热点,它在一定的应用背景下取得了很大的成功,如Cora和CiteSeer,但与此同时也存在很多尚待解决的问题,其中包括:第一,网页搜索没有能够充分利用搜索过程中网页与网页、网页与链接,以及链接与链接之间相互关联与约束的有关知识,因而无法更有效地提高搜索的效率和搜索的准确性。第二,网
【Abstract】 There are hidden and rich information for data mining in the topology of topic-specific websites. A new topic-specific association rules mining algorithm is proposed to further the research on this area- The key idea is to analyze the frequent hyperlinked relations between pages of different topics. In the topic-specific area, if pages of one topic are frequently hyperlinked by pages of another topic, we consider the two topics are relevant. Also, if pages of two different topics are frequently hyperlinked together by pages of the other topic, we consider the two topics are relevant. The initial experiments show that this algorithm performs quite well while guiding the topic-specific crawling agent and it can be applied to the further discovery and mining on the topic-specific website.
【Key words】 Association rule; Topic-specific crawling; Web mining;
- 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2003年10期
- 【分类号】TP393.09
- 【被引频次】2
- 【下载频次】69