节点文献
基于特征句抽取的网页去重研究
Research on Deletion of Feature Sentence Extraction Based Duplicated Web Pages
【Author】 Peng Yuan, Zhao Tiejun, Zheng Dequan, Yu Hao
【机构】 哈尔滨工业大学语言语音教育部-微软重点实验室;
【摘要】 去除重复网页一直是信息检索领域的一个待解决的问题。本文基于双语文章的内容,提出了一种抽取特征词和特征句,判别跨语言重复网页的方法。并将其运用到了跨语言的重复网页的识别上。实验结果表明:该方法对双语重复网页的识别准确率在86%以上,对单语重复网页的识别准确率在97.5%以上,达到了实用的程度,同时,该方法对于双语平行语料的自动挖掘也有一定的帮助。
【Abstract】 Deletion of duplicated web pages has been one of the problems that need to be solved in information retrieval. In this paper, according to the word frequency statistical theory, we put forward a duplicated web pages recognition algorithm by extracting feature words and feature sentences of web documents. We apply this approach in the recognition of cross language duplicated web pages. Experimental results show that this algorithm can reach a precision of 97.5% in mono-language deletion of duplicated web pages, and this algorithm can also reach a maximum precision of 86% when it is applied to deletion of duplicated web pages for cross language. This algorithm can also do some help to the automatic mining of parallel corpus from the internet.
【Key words】 Deletion of duplicated web pages; Feature word; Feature sentence; Cross-language;
- 【会议录名称】 全国第八届计算语言学联合学术会议(JSCL-2005)论文集
- 【会议名称】全国第八届计算语言学联合学术会议(JSCL-2005)
- 【会议时间】2005-08
- 【会议地点】中国南京
- 【分类号】TP391.1
- 【主办单位】南京师范大学、清华大学智能技术与系统国家重点实验室