节点文献

基于Fuzz与机器学习的二进制程序分析研究

Research on Binary Program Analysis Based on Fuzz and Machine Learning

【作者】 赵斌;

【导师】 胡鹏飞;

【作者基本信息】 山东大学 , 计算机技术(专业学位), 2024, 硕士

【摘要】 近几十年来,计算机技术飞速发展,其硬件和软件规模也发生了翻天覆地的变化。计算机已经承载了人类社会生活生产的方方面面,同时也引起了人们对程序安全问题的关注。目前市面上存在着大量无人维护且缺乏源代码支持的软件,其安全性令人担忧,而二进制程序分析作为计算机安全领域的关键研究方向,对于保障软件安全和防范潜在威胁具有重要意义。二进制程序安全分析通常被用来检测软件程序恶意行为,包括但不限于漏洞检测、恶意软件分类和逆向工程等场景。另外,该技术还被广泛用于生成具有全面路径覆盖的测试用例,包括符号执行、Fuzz等技术。总体而言,针对传统的二进制程序的分析技术通常可以划分为三类:基于静态的二进制分析技术、基于动态的二进制分析技术以及动静态混合的二进制分析技术。近年来,随着人工智能的快速发展,新的手段涌现,通过将机器学习、深度学习与二进制分析技术结合,为安全研究人员提供了更多工具和方法。而本文分别针对传统程序分析技术—Fuzz,和非传统提升手段—机器学习与程序分析相结合进行二进制函数识别,作了以下研究:(1)Fuzzing是经典的自动化漏洞分析工具,也称为模糊测试,它通过变异输入来测试程序,试图达到不同的代码分支,然而当测试的代码比较深、分支条件较为复杂时,其难以通过变异输入猜测到新的分支,分支探索被“卡住”。为解决此问题,本文选择性使用现代符号执行技术angr对“卡住”的分支输入建立符号约束进行求解,并且在模糊测试引擎转换到符号执行引擎时,提出了新的种子选择算法seeding和符号执行路径搜索算法Lofitracing,挑选高质量种子并避免符号执行追踪到一些内存不存在的地址,最终将解出的新输入送回到Fuzzing引擎中。实验表明该方法在真实软件二进制文件中识别的漏洞比目前几个先进方法识别的漏洞更多,代码覆盖率和检测效率也更高,在数据集中提高了 5%-10%,能够到达更深的代码分支,实现了更好的、更实用的漏洞检测。(2)函数识别是对恶意软件检测、常见漏洞检测和二进制工具等许多应用程序进行二进制程序分析的初步步骤。当前机器学习在函数识别领域应用较少,多数模型仅能识别函数的起始地址,未能准确识别整个函数范围。为解决此问题,本文提出一种新的方法,将函数识别任务转换成识别三种状态序列(指令包含状态,指令排除状态,函数结束状态)的任务,训练一个双向递归神经网络将状态序列与二进制程序相匹配,再将二进制文件输入到经过训练的模型中以输出其状态序列,最终将其序列进一步解码以计算出函数的起点、终点、边界以及范围。实验表明,本文提出的方法在二进制函数识别中的预测性能方面优于目前几个先进的方法,更是在函数范围识别这一最困难任务中表现突出。

【Abstract】 In recent decades,computer technology has developed rapidly,and the scale of its hardware and software has also undergone earth-shaking changes.Computers have carried all aspects of human social life and production,and have also attracted people’s attention to program security issues.Currently,there are a large number of unmaintained software on the market that lacks source code support,and its security is worrying.Binary program analysis,as a key research direction in the field of computer security,is of great significance for ensuring software security and preventing potential threats.Binary program security analysis is usually used to detect malicious behaviors of software programs,including but not limited to vulnerability detection,malware classification,reverse engineering and other scenarios.In addition,this technology is also widely used to generate test cases with comprehensive path coverage,including symbolic exccution,Fuzz and other technologies.In general,analysis technologies for traditional binary programs can usually be divided into three categories:static-based binary analysis technology,dynamic-based binary analysis technology,and dynamic-static hybrid binary analysis technology.In recent years,with the rapid development of artificial intelligence,new methods have emerged,combining machine learning,deep learning and binary analysis technology to provide security researchers with more tools and methods.This article conducts the following research on the traditional program analysis technology—Fuzz,and the non-traditional improvement method—the combination of machine learning and program analysis for binary function identification.(1)Fuzzing is a classic automated vulnerability analysis tool.It faces a common problem,that is,it tests programs by mutating inputs in an attempt to reach different code branches.However,when the tested code is relatively deep and the branch conditions are complex,it is difficult to guess new branches through mutation input.Branch exploration is"stuck".In order to solve this problem,this article selectively uses the modern symbolic execution technology angr to establish symbolic constraints for "stuck" branch inputs to solve,this paper selectively uses the modern symbolic exccution technology angr to solve the"stuck" branch input to establish symbolic constraints,and proposes a new seed selection algorithm seeding and symbolic execution path search algorithm Lofitracing when the fuzz testing engine is converted to a symbolic execution engine,select high-quality seeds and prevent symbolic execution from tracing to some addresses that do not exist in memory,and finally return the solved new input to the Fuzzing engine.The experiment shows that this method identifies more vulnerabilities in real software binary files than several advanced methods currently,and the code coverage is also higher.It improves by 5%-10%in the dataset and can achieve deeper code branches,achieving better and more practical vulnerability detection.(2)Function identification is a preliminary step in binary program analysis for many applications such as malware detection,common vulnerability detection,and binary tools.Currently,deep learning is rarely used in the field of function recognition.Most models can only identify the starting address of a function and fail to accurately identify the entire function range.In order to solve this problem,this paper proposes a new method to convert the function identification task into the task of identifying three state sequences(instruction inclusion state,instruction exclusion state,function end state),and train a bidirectional recurrent neural network to combine the state sequence with The binary program is matched,the binary file is input into the trained model to output its sequence of states,and finally its sequence is further decoded to calculate the start,end,boundary,and scope of the function.Experiments show that our proposed method outperforms current state-of-the-art methods in terms of predictive performance on real software datasets.It performs particularly well in the most difficult task of identifying function scope.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2025年 08期
  • 【分类号】TP309;TP181
节点文献中: 

本文链接的文献网络图示:

本文的引文网络