节点文献
HtmIParser提取网页信息的设计与实现
Design and Implementation of Web Information Extraction Based on HtmlParser
【摘要】 互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接、email和图片等,而HTML语言所表述的web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理,文中详细介绍了如何使用HtmlParser来提取网页当中的超链接信息,将其清洗后存入SQL数据库当中,以备后续工作使用。
【Abstract】 The rapid growth of the Web contents increases the need for some automatic tools to help people find the information among the magnanimous information sources such as tides,links,emails,pictures etc.The Web pages expressed by HTML,after analyzed by Internet Explorer,are only suitable for browse,but not for machine process- ing as the way of data exchange.The paper explains how to use HtmlParser to extract hyperlink information from web page,then store in SQL database after cleaning in information detail.
- 【文献出处】 江西理工大学学报 ,Journal of Jiangxi University of Science and Technology , 编辑部邮箱 ,2007年06期
- 【分类号】TP393.092
- 【被引频次】25
- 【下载频次】389