节点文献

基于结构与内容的网页主题信息提取研究

Structure and content-based extraction of topical information from Web pages

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 吴鹏飞孟祥增刘俊晓马凤娟

【Author】 WU Peng-fei,MENG Xiang-zheng,LIU Jun-xiao and MA Feng-juan(School of Communication,Shandong Normal Univ.,Jinan 250014,Shandong,China)

【机构】 山东师范大学传播学院山东师范大学传播学院 山东济南250014山东济南250014

【摘要】 结合HTML网页内部特征与外部的结构布局,提出采用映射表这种网页映射模式对网页视图进行变换,基于结构与启发式规则对网页进行区域分割与识别,并利用向量空间模型对网页内容分析,从而准确得到具有高语义内聚性的网页主题内容.实验结果表明,此方法对各种复杂结构的网页主题信息提取较为理想.

【Abstract】 Combining the Web page’s internal features and external structural layout,mapping table is suggested to tansform the view of Web page.The approach gets highly semantic cohesiveness of the topical contents of the Web page exactly,based on the structure and revelatory rules for Web page’s segmentation and identification and the use of the vector space model for Web(content) analysis.Experimental results show that this method is more ideal for the topical information extraction of complexstructure Web pages.

【基金】 山东省自然科学基金资助项目(Y2005G21)
  • 【文献出处】 山东大学学报(理学版) ,Journal of Shandong University(Natural Science) , 编辑部邮箱 ,2006年03期
  • 【分类号】TP393.092;TP391.1
  • 【被引频次】43
  • 【下载频次】472
节点文献中: