节点文献

基于深度学习和风格迁移神经网络的服装设计优化

Garment design optimization based on deep learning and style migration neural network

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 俞博涵; 宛俊勇;

【Author】 YU Bohan;WAN Junyong;School of Art, Anhui Yangzi Vocational and Technical College;School of Fine Arts, Anhui Normal University;

【机构】 安徽扬子职业技术学院艺术学院; 安徽师范大学美术学院;

【摘要】 为解决传统风格迁移方法在图像结构保持与风格表达之间难以兼顾的问题,提出了一种融合坐标注意力机制与Transformer结构的风格迁移神经网络模型CATrans-VGG19。首先采用VGG19作为主干网络对服装图像的多尺度结构特征进行深度编码,随后引入坐标注意力机制,嵌入空间位置信息,强化对关键纹理与局部结构的响应,最后构建基于Transformer的编码-解码融合模块,利用多头注意力实现内容与风格特征的全局建模与深度交互,增强风格表达与结构还原的协同性。为验证模型性能,基于MS-COCO与WikiArt数据集开展实验,并从图像质量、结构一致性和计算效率等维度与VGG19、AAMS和ArtFlow模型进行对比,重点测试峰值信噪比、结构相似性、迁移效果、推理时间及显存占用等核心指标。结果表明,该模型结构相似性指标达到0.743,峰值信噪比达到29.8 dB,推理时间仅为15 ms,占用显存2.1 GB,兼顾了迁移质量与计算效率。

【Abstract】 To address the challenge of balancing image structure preservation and style expression in traditional style transfer methods, a style transfer neural network model CATrans-VGG19 that integrates coordinate attention mechanism and Transformer structure is proposed. Firstly, VGG19 is used as the backbone network to deeply encode the multi-scale structural features of clothing images. Subsequently, a coordinate attention mechanism is introduced to enhance the response to key textures and local structures by embedding spatial position information. Finally, a Transformer based encoding decoding fusion module is constructed, utilizing multi head attention to achieve global modeling and deep interaction of content and style features, enhancing the synergy between style expression and structural restoration. To verify the performance of the model, experiments were conducted based on the MS-COCO and WikiArt datasets, and compared with three models, VGG19, AAMS, and ArtFlow, in terms of image quality, structural consistency, and computational efficiency. The focus is on testing core indicators such as peak signal-to-noise ratio, structural similarity, transfer effect, inference time, and video memory usage. The experimental results show that the model achieves a structural similarity index of 0.743, a peak signal-to-noise ratio of 29.8 dB, a inference time of only 15 ms, and a low video memory occupation of 2.1 GB, combining transfer quality and computational efficiency.

【基金】 安徽省重点研究与开发计划项目(202009A251347RYZ)
  • 【文献出处】 河南工程学院学报(自然科学版) ,Journal of Henan University of Engineering(Natural Science Edition) , 编辑部邮箱 ,2026年01期
  • 【分类号】TP18;TP391.41;TS941.2
  • 【下载频次】39
节点文献中: 

本文链接的文献网络图示:

本文的引文网络