节点文献
Web信息抽取技术综述
Survey of Web information extraction technologies
【摘要】 快速高效地获取网页主题信息的需求使得Web信息抽取技术成为信息技术领域的研究热点。现有的Web信息抽取技术大致可以归纳为基于统计理论的、基于视觉特征的、基于DOM树结构的和基于模板的几类。由于网页文本本身具有树结构并且具有一定的相似性,基于DOM树结构和基于模板的抽取技术发展很快而且已经得到了广泛的应用。分别论述了上述几类技术在近几年来的研究进展,从自动化程度、适用范围和复杂性三个角度分析对比了几类技术的优缺点。
【Abstract】 Web information extraction technology has been made the focus of the field of information technology by the needs of obtaining the topic contents of Web pages more efficiently. Existing technologies of this field could be classified into the following four categories,statistics based technology,vision based technology,DOM tree based technology and template based technology. The DOM tree based technology and template based technology had gained a rapid development and a wide employment because of the special structure and similarity owned by Web pages. This paper made a detailed survey and analysis of the above four technologies as well as the comparison of their advantages and disadvantages from points of automation,application filed and complexity.
【Key words】 Web information extraction; Web page noise; URL clustering; DSE algorithm; RoadRunner system; MDR algorithm; vision feature; template;
- 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2010年12期
- 【分类号】TP393.09
- 【被引频次】116
- 【下载频次】1585