节点文献
网页结构模板生成新方法研究
New method of generating template from webpage structure
【摘要】 Web页面所表达的主要信息通常隐藏在大量无关的结构和文字中,使用户不能迅速获取主题信息,限制了Web的可用性。为了高效地抽取基于模板的网页主题信息,提出了一种新的从HTML网页结构分析入手的模板生成方法。该方法以文档对象模型(DOM)为基础,通过对网页对应的DOM树层次结构进行分析,来判断两个网页是否相似,结构上相似的网页可以作为一个样本集。利用生成的样本集可以比较方便的抽象出网页结构模板,实现高效的信息抽取。实验表明,该方法准确率可达97%。
【Abstract】 Web is a vast resource of information,but the main information on a web page is always hidden among unimportant features such as unnecessary images and extraneous links,which makes it difficult for the users to acquire the topical information.In order to automatically extract topical information from template-based web pages efficiently,a new template-generating method based on the structure analysis of HTML webpages is proposed in this paper.On the basis of document object model(DOM),the similarity of two pages can be calculated by analyzing their DOM tree hierarchy structure,then the similar pages in structure are put into a sample-collection,with which the structure-template of pages can be deduced with little effort.In this way,information from webpages can be extracted efficiently.The experiments indicate that the accuracy of this new method reaches 97%.
【Key words】 DOM; structure analysis; webpage similarity; sample collection; template;
- 【文献出处】 北京机械工业学院学报 ,Journal of Beijing Institute of Machinery , 编辑部邮箱 ,2007年03期
- 【分类号】TP393.092
- 【被引频次】10
- 【下载频次】175