节点文献

HtmIParser提取网页信息的设计与实现

Design and Implementation of Web Information Extraction Based on HtmlParser

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄颖; 黄治平;

【Author】 HUANG Ying HUANG Zhi-ping 1.Facuhy of Information Engineering,Jiangxi University of Science and Technology,Ganzhou 341000,China 2.Gannan Teachers College,Ganzhou 341000,China

【机构】 江西理工大学信息工程学院; 赣南师范学院 江西 赣州 341000; 江西 赣州 341000;

【摘要】 互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接、email和图片等,而HTML语言所表述的web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理,文中详细介绍了如何使用HtmlParser来提取网页当中的超链接信息,将其清洗后存入SQL数据库当中,以备后续工作使用。

【Abstract】 The rapid growth of the Web contents increases the need for some automatic tools to help people find the information among the magnanimous information sources such as tides,links,emails,pictures etc.The Web pages expressed by HTML,after analyzed by Internet Explorer,are only suitable for browse,but not for machine process- ing as the way of data exchange.The paper explains how to use HtmlParser to extract hyperlink information from web page,then store in SQL database after cleaning in information detail.

【关键词】 HtmIParser; 信息提取; 网页解析;
【Key words】 htmlparser; information extraction; web analysis;
  • 【文献出处】 江西理工大学学报 ,Journal of Jiangxi University of Science and Technology , 编辑部邮箱 ,2007年06期
  • 【分类号】TP393.092
  • 【被引频次】25
  • 【下载频次】389
节点文献中: 

本文链接的文献网络图示:

本文的引文网络