节点文献

网页结构模板生成新方法研究

New method of generating template from webpage structure

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 冯少卿; 都云程;

【Author】 FENG Shao-qing,DU Yun-cheng(Chinese Information Processing Research Center,Beijing Information Technology Institute,Beijing 100101,China)

【机构】 北京信息工程学院中文信息处理研究中心; 北京信息工程学院中文信息处理研究中心 北京100101; 北京100101;

【摘要】 Web页面所表达的主要信息通常隐藏在大量无关的结构和文字中,使用户不能迅速获取主题信息,限制了Web的可用性。为了高效地抽取基于模板的网页主题信息,提出了一种新的从HTML网页结构分析入手的模板生成方法。该方法以文档对象模型(DOM)为基础,通过对网页对应的DOM树层次结构进行分析,来判断两个网页是否相似,结构上相似的网页可以作为一个样本集。利用生成的样本集可以比较方便的抽象出网页结构模板,实现高效的信息抽取。实验表明,该方法准确率可达97%。

【Abstract】 Web is a vast resource of information,but the main information on a web page is always hidden among unimportant features such as unnecessary images and extraneous links,which makes it difficult for the users to acquire the topical information.In order to automatically extract topical information from template-based web pages efficiently,a new template-generating method based on the structure analysis of HTML webpages is proposed in this paper.On the basis of document object model(DOM),the similarity of two pages can be calculated by analyzing their DOM tree hierarchy structure,then the similar pages in structure are put into a sample-collection,with which the structure-template of pages can be deduced with little effort.In this way,information from webpages can be extracted efficiently.The experiments indicate that the accuracy of this new method reaches 97%.

【关键词】 DOM; 结构分析; 网页相似; 样本集; 模板;
【Key words】 DOM; structure analysis; webpage similarity; sample collection; template;
【基金】 863计划重点项目(2006AA010105);北京市属市管高校人才强教计划项目(PXM2007-014224-044677,PXM2007-014224-044676);北京市教委科技发展计划项目(KM200710772010)
  • 【文献出处】 北京机械工业学院学报 ,Journal of Beijing Institute of Machinery , 编辑部邮箱 ,2007年03期
  • 【分类号】TP393.092
  • 【被引频次】10
  • 【下载频次】175
节点文献中: 

本文链接的文献网络图示:

本文的引文网络