节点文献

基于深度学习的自然语言生成SQL方法研究与应用

Research and Application on Method of Generating SQL through Natural Language Based on Deep Learning

【作者】 葛岩;

【导师】 白天;

【作者基本信息】 吉林大学 , 软件工程(专业学位), 2021, 硕士

【摘要】 随着计算机技术的不断发展,人类社会中的各个方面开始与之产生紧密联系。人们日常生产生活中产生的海量数据大多以电子化的形式存储在关系型数据库中,在对这些数据进行访问时,往往需要编写SQL(Structured Query Language)来对数据库进行操作。但是,SQL本质上是一种计算机编程语言,编写SQL需要一定的专业知识,此外,还需要了解所访问的数据库模式。通过自然语言来与数据库进行交互查询数据,既能节省用户学习专业知识和了解数据库模式的时间,又能提高查询的效率。因此,如何根据自然语言生成SQL来查询数据库,具有很重要的研究与应用价值,并逐渐成为自然语言处理领域的热点之一。近年来,随着深度学习技术的成熟,越来越多的研究者开始将深度学习技术应用到自然语言生成SQL(NL2SQL)任务中。虽然目前深度学习模型在该领域已取得了较好的效果,但是仍存在一些待解决的问题。如在将自然语言向SQL转换时,往往需要识别出自然语言中的所提到的数据库表和列,这就要求在对数据库模式(数据库中的表和列以及它们之间的关系)进行编码时,需要将数据库中的模式信息详尽的表示出来。另外,在中文自然语言生成SQL任务中,还存在着自然语言查询中的描述和数据库存储的数据表述不一致的问题。针对这些问题,本文从提升深度学习模型生成SQL的准确率的角度出发,基于目前在该任务上较为先进的IRNet模型,做出了以下工作与贡献:(1)针对英文NL2SQL问题,在IRNet模型中加入了门控图神经网络(GGNN)用于编码数据库模式,将数据库模式中的全局信息融入到每个表名和列名的词嵌入中,使得模型在生成SQL中的数据库表名和列名时能够感知到更多的数据库模式的上下文关系,提升了模型生成SQL的准确匹配率。(2)针对英文NL2SQL问题,将数据库的值引入到IRNet模型中。通过计算自然语言查询和数据库值之间的注意力,从而对数据库值以及相应的自然语言查询进行匹配,将数据库值与列之间的关联关系引入模型,使其更准确地预测SQL中的列名,提升了模型生成SQL的准确匹配率。(3)针对中文NL2SQL问题,在IRNet的基础上,加入了预训练语言模型,使得IRNet可以处理中文自然语言生成SQL语句的问题。本文采用跨语言预训练模型来对数据库模式和自然语言查询进行编码,这样就在一定程度上解决了中文自然语言与英文数据库模式之间的映射问题。(4)针对中文自然语言查询数据库需求,在改进的深度自然语言生成SQL语句模型之上,开发了一个自然语言数据库查询系统。在该系统中,用户只要选择需要查询的数据库并输入相应的自然语言查询,系统会自动将自然语言查询转换为SQL语句并在相应的数据库上执行,并将结果反馈给用户。

【Abstract】 With the continuous development of computer technology,all aspects of human society have become closely connected with it.Most of the massive data generated in people’s daily production and life are stored in relational databases in electronic form.When accessing these data,it is often necessary to write SQL(Structured Query Language)to operate the database.However,SQL is essentially a computer programming language and writing SQL requires a certain level of expertise,in addition to an understanding of the database schema being accessed.Using natural language to interact with the database to query data can not only save users the time to learn professional knowledge and understand the database schema,but also improve the efficiency of the query.Therefore,how to generate SQL based on natural language to query the database has very important research and application value,and has gradually become one of the hot spots in the field of natural language processing.In recent years,with the success of deep learning techniques,more and more researchers have started to apply deep learning techniques to NL2SQL(Natural Language to Structured Query Language)tasks.Although deep learning techniques have achieved good results in this field,there are still some problems to be solved.For example,when converting natural language to SQL,it is often necessary to identify the database tables and columns mentioned in the natural language,which requires an exhaustive representation of the database schema(the tables and columns in the database and the relationships between them)when encoding the database schema.Moreover,in the task of generating SQL in Chinese natural language,there is a problem that the description in the natural language query is inconsistent with the data stored in the database.To address these issues,this paper makes the following work and contributions from the perspective of improving the accuracy of SQL generation by deep learning models,based on the IRNet model,which are:(1)In this paper,a gated graph neural network(GGNN)is added to the IRNet model for encoding database schema,incorporating global information in the database schema into the word embedding of each table name and column name,enabling the model to perceive more contextual relationships of the database structure when generating database table names and column names in SQL.(2)In this paper,the association relationship between database content and columns is introduced into the model by calculating the attention score between natural language queries and database content so as to match database content as well as the corresponding natural language queries and make it better predict the column names in SQL.(3)In this paper,a pre-trained language model is added to IRNet to handle the problem of generating SQL through Chinese natural language.A cross-lingual pre-training model is used to encode database schemas and natural language,so that the mapping problem between Chinese natural language and English database schemas can be solved to some extent.(4)In this paper,a natural language database query system is developed on top of the enhanced deep learning model.In this system,the user simply selects the database to be queried and enters the corresponding natural language query question,and the system automatically converts the natural language question into a SQL statement and executes it on the corresponding database,then gives the result back to the user.

  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2022年 01期
  • 【分类号】TP18;TP311.13
  • 【被引频次】1
  • 【下载频次】201
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络