节点文献
基于块分布的新闻网页内容提取
News content extraction based on block distribution
【摘要】 提出一种新的新闻网页内容提取方法。与已有的研究相比,它自动判别网页是否含有主内容,并且回避了模板和DOM-Tree方法所带来的局限。主要工作包括:①提出了一种网页分块方法,通过一趟遍历将网页主内容和噪声划分到不同的块中;②提出网页块分布的概念并研究了块分布的属性,根据块分布可以有效地使用分类方法来判别网页是否有主内容,采用孤立点分析的方法从网页块分布中提取主内容。本文通过理论和实验证明了该方法的有效性。
【Abstract】 An approach to extract news contents automatically from news web pages is proposed.Compared with existing methods,this approach can determine whether a web page contains news content first,then extract the news contents without using DOM-Tree or template.A new concept of Block is introduced and by one traversal the approach partitions web page into main content block and noise block.Further more,the concept of Web Page Block Distribution is introduced and the features of Block Distribution are investigated.The use of Block Distribution can effectively determine whether a web page contains news contents.Experiments show the approach is effective in extraction of news contents.
【Key words】 computer application; Web contents extracting; block distribution; Web mining;
- 【文献出处】 吉林大学学报(工学版) ,Journal of Jilin University(Engineering and Technology Edition) , 编辑部邮箱 ,2009年05期
- 【分类号】TP393.092
- 【被引频次】11
- 【下载频次】248