节点文献

基于参数高效扩散微调的少样本参考图驱动图像补全(英文)

Few-shot exemplar-driven inpainting with parameter-efficient diffusion fine-tuning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 杨诗远; 顾峥; 郝文月; 汪毅; 蔡怀宇; 陈晓冬;

【Author】 Shiyuan YANG;Zheng GU;Wenyue HAO;Yi WANG;Huaiyu CAI;Xiaodong CHEN;Key Laboratory of Optoelectronics Information Technology, Ministry of Education, School of Precision Instruments and Optoelectronic Engineering, Tianjin University;State Key Laboratory for Novel Software Technology, Nanjing University;

【通讯作者】 陈晓冬;

【机构】 天津大学精密仪器与光电子工程学院光电信息技术教育部重点实验室; 南京大学计算机软件新技术国家重点实验室;

【摘要】 文本到图像的扩散模型在图像生成方面展现了卓越的能力,并已广泛应用于图像补全任务。尽管文本提示能够为有条件的图像补全提供直观指导,但用户往往希望通过提供参考图像为特定对象补全个性化外观。然而,现有的参考图驱动图像补全方法难以实现高保真度的补全效果。为解决这一问题,我们基于预训练的文本驱动图像补全模型提出一种即插即用的低秩适配(Lo RA)模块。该模块通过少样本微调学习参考图像的特定特征,显著提升了对自定义参考图像的拟合能力,并且无需在大规模数据集上进行大量训练。此外,引入GPT-4V提示词和先验噪声初始化技术,进一步提升补全结果的保真度。简而言之,去噪扩散过程首先从由复合参考—背景图像派生的初始噪声开始,进而由GPT-4V从参考图中生成的丰富提示词引导后续生成过程。大量实验表明,我们的方法在定性和定量指标上都达到目前最高水平,为用户提供了一个具有更强定制化能力的参考图驱动图像补全工具。

【Abstract】 Text-to-image diffusion models have demonstrated impressive capabilities in image generation and have been effectively applied to image inpainting. While text prompt provides an intuitive guidance for conditional inpainting, users often seek the ability to inpaint a specific object with customized appearance by providing an exemplar image. Unfortunately, existing methods struggle to achieve high fidelity in exemplar-driven inpainting.To address this, we use a plug-and-play low-rank adaptation(LoRA) module based on a pretrained text-driven inpainting model. The LoRA module is dedicated to learn the exemplar-specific concepts through few-shot fine-tuning,bringing improved fitting capability to customized exemplar images, without intensive training on large-scale datasets.Additionally, we introduce GPT-4V prompting and prior noise initialization techniques to further facilitate the fidelity in inpainting results. In brief, the denoising diffusion process first starts with the noise derived from a composite exemplar–background image, and is subsequently guided by an expressive prompt generated from the exemplar using the GPT-4V model. Extensive experiments demonstrate that our method achieves state-of-the-art performance,qualitatively and quantitatively, offering users an exemplar-driven inpainting tool with enhanced customization capability.

【基金】 Project supported by the National Natural Science Foundation of China (No.82027801)
  • 【文献出处】 Frontiers of Information Technology & Electronic Engineering ,信息与电子工程前沿(英文) , 编辑部邮箱 ,2025年08期
  • 【分类号】TP391.41
  • 【下载频次】11
节点文献中: 

本文链接的文献网络图示:

本文的引文网络