节点文献
文档复制检测技术
Document copy detection technology
【摘要】 随着数字图书馆和互联网的飞速发展,数字化文档唾手可得。近年来学术剽窃现象屡见报端,互联网上日益增多的重复网页降低了检索效率,给用户带来不便。文档复制检测技术在保护知识产权和优化搜索引擎方面起着重要作用,是近年来数据库安全领域研究的热点。文档复制检测方法有两类:一是基于词频统计的方法,一是基于字符串匹配的方法。本文详尽分析了现有基于这两类方法的复制检测技术,并指出它们的优缺点,针对两类方法都存在的问题提出一些改进方案。最后总结了复制检测技术应满足的特性,讨论了检测方法的准确性和文档分解规则。
【Abstract】 With the rapid development of digital library and the intemet,digital documents are easily acquired.Recent years,there are many news about plagiarism on reseach,and the number of duplicated pages on the web is increasing,which lower the efficiency of search and put users to inconvenience.The copy detection technique plays an important role on intellectual property protection and information retrieval,this technique is the hot topic in field of database security.There are two approachs to copy detection, one is based on the frequency of words which appear in the document,the other is based on the string match.The exsited systems based on the two approachs are analyzed and the merit and shortage is pointed out.The improved scheme is proposed.The characters that copy detection technique should satisfy are summarized,and the veracity of detection and rules of document division are di- scussed.
- 【文献出处】 燕山大学学报 ,Journal of Yanshan University , 编辑部邮箱 ,2007年05期
- 【分类号】TP391.1
- 【被引频次】19
- 【下载频次】296