节点文献

TECCD:基于树嵌入的代码克隆检测方法

TECCD: A Tree Embedding Approach for Code Clone Detection

【作者】 高毅;

【导师】 王赞; 孟少卿;

【作者基本信息】 天津大学 , 软件工程, 2019, 硕士

【摘要】 在软件工程领域,学者们对代码克隆检测技术的研究从来未停止过。代码克隆检测的目的是为了找出软件系统中存在的克隆,通过分析克隆对软件质量的影响来利用有益克隆,同时对软件质量造成威胁的有害克隆进行规避或重构,从而提升软件系统的质量,提高软件人员的开发效率并减少软件维护成本。到目前为止,在克隆检测领域不同方法和技术的积累,主流检测技术分为基于代码语法结构的检测和非语法结构检测。基于非语法结构检测又分为基于文本检测,基于令牌检测和基于代码度量值检测等,这类方法的优点是检测速度快,然而由于未考虑代码结构之间的相似性,这类技术普遍对高级别克隆代码检测的效果不甚理想。反之,基于语法结构检测的技术由于将代码处理为解析树或程序依赖图等,它们将代码之间的结构信息用于最后的检测过程,因此对高级别克隆代码表现出更好的检测效果,但与此同时,基于语法结构检测的技术存在最大的问题是检测效率低下,这是因为将代码转换为树结构或图结构,以及在这些结构上使用到的匹配算法开销比较大,鉴于此,这类技术的应用同样受到了很大的限制。近年来,随着深度学习技术不断提高代码的表达能力,并提高了代码克隆检测的技术水平,这些方法通常需要从抽象语法树向二叉树的转换来合并代码的语法信息,而这样会造成信息损失和额外开销。此外,这些方法使用术语项嵌入技术,而这往往需要大量的训练数据集。为了解决传统检测技术存在的问题,以及近来基于深度学习克隆检测方法的不足之处,本文介绍了一种基于树嵌入技术来进行代码克隆检测,我们的方法首先运用树嵌入技术以获得代码抽象语法树中每个中间层的节点向量,该节点向量捕获了语法树的结构信息。然后,我们使用轻量级方法从相关节点向量中合成一个树向量。最后,通过计算树向量之间的欧几里得距离,以确定代码克隆。在我们称为TECCD的工具中的应用,使用Big Clone Bench(BCB)和6个其他规模的开源Java项目进行评估。结果表明,我们的方法具有良好的精度和召回率,并优于现有的方法。

【Abstract】 In the field of software engineering,scholars have never stopped the research of code clone detection technology.The purpose of code clone detection is to find out the clones existing in software system,make use of them scientifically by analyzing the impact of cloning on software quality,and reconstruct or avoid harmful clones that threaten software quality,so as to improve the quality of software system,improve the development efficiency of software personnel and reduce the maintenance cost of software system.Up to now,in the field of clone detection,different methods and techniques have been accumulated.Mainstream detection technologies include code-based grammar structure detection and Non-syntactic structure detection.Non-syntactic structure detection is divided into text-based detection,token-based detection and code Metrics-based detection.The advantages of these methods are f AST detection speed.However,due to the lack of consideration of the similarity between code structures,the effect of this kind of technology on high-level cloned code detection is generally unsatisfactory.Conversely,the technology based on grammatical structure detection has better detection effect for high-level cloned code because it processes code into parse tree or program dependency graph,which uses structural information between codes for final detection.But at the same time,the biggest problem of the technology based on grammatical structure detection is the inefficiency of detection,which is due to the conversion of code to program dependency graph.Tree structure or graph structure,as well as matching algorithms used in these structures,are expensive.In view of this,the application of this kind of technology is also greatly limited.Recently,deep learning techniques has been adopted to improve the code representation capability,and improve the state-of-the-art in code clone detection.These approaches usually require a transformation from AST to binary tree to incorporate syntactical information,which introduces overheads.Moreover,these approaches conduct term-embedding,which requires large training datasets.In this paper,we introduce a tree embedding technique to conduct clone detection.Our approach first conducts tree embedding to obtain a node vector for each intermediate node in the AST,which captures the structure information of ASTs.Then we compose a tree vector from its involving node vectors using a lightweight method.Lastly Euclidean distances between tree vectors are measured to determine code clones.We implement our approach in a tool called TECCD and conduct an evaluation using the Big Clone Bench(BCB)and 6 other large scale Java projects.The results show that our approach achieves good accuracy and recall and outperforms existing approaches.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2022年 01期
节点文献中: