节点文献
基于网页结构挖掘的信息提取
Extracting Information by Mining Structures of Web Pages
【摘要】 本文提出了两种细粒度的、基于网页结构挖掘的信息提取方法,比较了它们的优缺点,并给出了相应具体实现的性能测试和结果分析。
【Abstract】 To simplify the task of obtaining information from the vast number of information sources that are available on the WWW,we have developed two different methods to extract information of fine grain.This paper firstly de- scribes the principles of the two methods,which work by mining structures of Web pages,and then compares the ad- vantages and disadvantages of them.Finally,we test the performance of the two methods and analyze the experiment results.
【关键词】 信息提取;
网页结构挖掘;
重复模式;
时间特征;
RSS;
【Key words】 Information extraction; Mining structures of Web pages; Repeated pattern; Time characteristic; RSS;
【Key words】 Information extraction; Mining structures of Web pages; Repeated pattern; Time characteristic; RSS;
- 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2006年03期
- 【分类号】TP311.13
- 【被引频次】12
- 【下载频次】256