节点文献

基于生成对抗网络的文本描述生成图像研究与软件开发

Research and Software Development of Text Description Generating Image Based on Generative Adversarial Network

【作者】 王博文;

【导师】 钟玲;

【作者基本信息】 沈阳工业大学 , 软件工程, 2023, 硕士

【摘要】 文本描述生成图像是人工智能领域应用较为广泛的研究方向之一,它可以根据输入的文字描述自动生成相应的图像信息。文本描述生成图像通过应用深度学习技术,取得了良好的研究成果,具有较高的应用价值。目前文本描述生成图像在少量训练数据下的图像生成精度及在自然场景图像生成质量等方面仍存在较大的提升空间。针对上述问题,本文提出了基于场景图法文本描述生成图像模型的改进模型ASG-GAN(Advanced Scene Graph-GAN),模型的改进方向主要为模型的结构优化与文本的结构化表达,通过模型结构的优化,有利于模型在少量训练数据下的生成图像质量及精度的提升,而描述文本的结构化表达,有利于模型在自然场景下提升图像的生成质量。本文提出的ASG-GAN模型将堆叠式生成对抗网络引入到了基于场景图法的文本描述生成图像模型中,并通过引入Scene Graph Parser,将原有的场景图输入转变为描述文本的直接输入,增强模型的普遍适用性。在增加的第二层生成对抗网络模型中引入预训练的文本编码器获取原始描述文本的文本嵌入向量并通过条件增强技术生成条件向量,解决长文本嵌入向量条件流性不连续的问题。将第一层网络生成的图像与条件向量连接,传递至深层卷积的残差网络生成器中学习文本与图像的多峰向量特征,并输出生成图像。通过深层的可感知匹配鉴别器,鉴别生成图像与真实图像数据,形成完整的生成对抗体系。最终,通过多周期的迭代训练,实现改进模型。为了方便模型的便捷使用及二次训练,本文设计并实现了基于文本描述生成图像软件系统,该系统基于Python语言开发,前端框架为Tkinter。系统实现了模型的图像生成功能以及模型的可视化训练功能,在可视化训练中可自定义训练属性及训练数据集。使用COCO数据集以Inception Score、Fréchet Inception Distance作为客观评价标准,进行实验验证。实验表明,对比原有模型ASG-GAN在Inception Score评价体系中生成图像质量评分提升了7.8%,对比其他模型,在Fréchet Inception Distance评价体系中,生成图像质量评分均高于对比模型。ASG-GAN模型有效地提升了原有模型的生成图像分辨率,并提高了图像的生成质量,丰富了生成图像的细节,对比其他模型具有更强的多物体复杂文本及自然场景适应能力。

【Abstract】 Image generation from text description is one of the most widely used research directions in the field of artificial intelligence,which can automatically generate the corresponding image information according to the input text description.Through the application of deep learning technology,the image generated by text description has obtained good research results and has high application value.At present,there is still a large room for improvement in the accuracy of image generation and the quality of natural scene image generation with a small amount of training data.To solve the above problems,this thesis proposes an improved model ASG-GAN(Advanced Scene Graph-GAN),which generates image models based on the text description of scene diagrams.The improvement direction of the model is mainly model structure optimization and text structured expression.It is beneficial to improve the image quality and accuracy of the model under a small amount of training data,and the structured expression of the description text is conducive to improving the image generation quality of the model under natural scenes.The ASG-GAN model proposed in this thesis introduces the stacked generative adversarial network into the text description generating image model based on Scene Graph method,and converts the original scene graph input into the direct input of description text by introducing Scene Graph Parser to enhance the universal applicability of the model.A pre-trained text encoder is introduced into the second layer generative adversarial network model to obtain the text embedding vector of the original description text and generate the conditional vector by condition enhancement technology,which solves the problem of conditional flow discontinuity of long text embedding vector.The image generated by the first layer network is connected with the conditional vector,transferred to the deep convolutional residual network generator to learn the multi-modal vector features of text and image,and output the generated image.Through the deep perceptive matching discriminator,the generated image and the real image data are identified to form a complete generative antagonism system.Finally,the improved model is realized through multi-cycle iterative training.In order to facilitate the convenient use and secondary training of the model,this thesis designs and implements an image generation software system based on text description.The system is developed based on Python language and the front-end framework is Tkinter.The system realizes the image generation function of the model and the visual training function of the model.In the visual training,the training attributes and training data set can be customized.Using COCO data set,Inception Score and Frechet Inception Distance were used as objective evaluation criteria for experimental verification.Experiments show that compared with the original model ASG-GAN,the generated image quality Score in the Inception Score evaluation system is improved by 7.8%;Compared with other models,the generated image quality score in the Frechet Inception Distance evaluation system is higher than that of the comparison model.The ASG-GAN model effectively improves the image resolution of the original model,improves the image quality,and enrichis the details of the generated image.Compared with other models,it has stronger adaptability to complex text and natural scenes with multiple objects.

  • 【分类号】TP391.41
节点文献中: 

本文链接的文献网络图示:

本文的引文网络