节点文献

基于特征句抽取的网页去重研究

Research on Deletion of Feature Sentence Extraction Based Duplicated Web Pages

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 彭渊赵铁军郑德权于浩

【Author】 Peng Yuan, Zhao Tiejun, Zheng Dequan, Yu Hao

【机构】 哈尔滨工业大学语言语音教育部-微软重点实验室

【摘要】 去除重复网页一直是信息检索领域的一个待解决的问题。本文基于双语文章的内容,提出了一种抽取特征词和特征句,判别跨语言重复网页的方法。并将其运用到了跨语言的重复网页的识别上。实验结果表明:该方法对双语重复网页的识别准确率在86%以上,对单语重复网页的识别准确率在97.5%以上,达到了实用的程度,同时,该方法对于双语平行语料的自动挖掘也有一定的帮助。

【Abstract】 Deletion of duplicated web pages has been one of the problems that need to be solved in information retrieval. In this paper, according to the word frequency statistical theory, we put forward a duplicated web pages recognition algorithm by extracting feature words and feature sentences of web documents. We apply this approach in the recognition of cross language duplicated web pages. Experimental results show that this algorithm can reach a precision of 97.5% in mono-language deletion of duplicated web pages, and this algorithm can also reach a maximum precision of 86% when it is applied to deletion of duplicated web pages for cross language. This algorithm can also do some help to the automatic mining of parallel corpus from the internet.

【基金】 本文承国家自然科学基金(No:60302021);黑龙江省自然科学基金(No:F2004-04)的资助
  • 【会议录名称】 全国第八届计算语言学联合学术会议(JSCL-2005)论文集
  • 【会议名称】全国第八届计算语言学联合学术会议(JSCL-2005)
  • 【会议时间】2005-08
  • 【会议地点】中国南京
  • 【分类号】TP391.1
  • 【主办单位】南京师范大学、清华大学智能技术与系统国家重点实验室
节点文献中: 

本文链接的文献网络图示:

本文的引文网络