节点文献

基于Low-IDF-SIG的句子重复检测

Sentence Near-Duplicate Detection Based on Low-IDF-SIG

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 俞昊旻张玥张奇黄萱菁

【Author】 Haomin Yu,Yue Zhang,Qi Zhang,XuanJing Huang Fudan University,School of Computer Science and Technology,Shanghai,201203

【机构】 复旦大学计算机科学与技术学院

【摘要】 随着互联网上数据的爆炸式的增长,互联网上产生了大量的重复数据。这些重复数据给搜索引擎、观点挖掘等许多Web应用带来了严峻的问题。目前绝大部分的拷贝检测的算法均着重考虑文档级别,这些方法不能有效地检测出两个文档中只有一部分互为拷贝的情况。而句子级别的拷贝检测正是解决这类问题的一个必要步骤。本文提出了一种有效并且快速的句子级别的特征抽取方法——Low-IDF-Sig算法,并基于该算法实现了一个可以高效地找出句子级别拷贝的检测系统。为了对本文提出的方法的精度及效率进行评测,我们还在一个真实的语料库上对提出的方法与其他方法进行了比较。实验结果证明本文提出的方法能有效地提高句子级别的拷贝检测任务的效率和精度。

【Abstract】 Because of the explosion of the Internet,enormous duplicated data cause serious problem for search engine,opinion mining and many other Web applications.Most existed near-duplicate detection approaches focus on document level,so these approaches are not able to find out the duplicated part that is just a small piece of two documents.And near-duplicate detection on sentence level is a key step to solve such problem.An effective and efficient feature extraction algorithm—Low-IDF-Sig algorithm was proposed by this paper,and an efficient near-duplicate detection system on sentence level was built based on this algorithm.For evaluation,the proposed method was compared with other approaches on a real corpus.Experimental results show that our proposed method can improve both precision and efficiency of near-duplicate detection of sentence level.

【关键词】 句子级别拷贝检测
【Key words】 Sentence LevelNear Duplicate Detection.
  • 【会议录名称】 第六届全国信息检索学术会议论文集
  • 【会议名称】第六届全国信息检索学术会议
  • 【会议时间】2010-08-12
  • 【会议地点】中国黑龙江牡丹江
  • 【分类号】TP391.1
  • 【主办单位】中国中文信息学会信息检索与内容安全专业委员会
节点文献中: