节点文献

面向生成式视觉感知的细粒度直接偏好对齐框架

Fine-Grained Direct Preference Alignment Framework for Generative Visual Perception

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 谢涛袁玉轩左旺孟李瑞峰赵立军

【Author】 XIE Tao;YUAN Yuxuan;ZUO Wangmeng;LI Ruifeng;ZHAO Lijun;School of Mechatronics Engineering,Harbin Institute of Technology;Faculty of Computing,Harbin Institute of Technology;

【通讯作者】 赵立军;

【机构】 哈尔滨工业大学机电工程学院哈尔滨工业大学计算学部

【摘要】 基于多模态大语言模型的生成式指代分割方法缺乏对生成质量提升途径的深层探索,受限于监督微调的模仿机制,在复杂场景中面临语义定位偏差与掩码边界粗糙的挑战.为此,文中提出面向生成式视觉感知的细粒度直接偏好对齐框架(Fine-Grained Direct Preference Alignment Method for Generative Visual Perception, FG-DPA),将DPO(Direct Preference Optimization)从文本领域迁移至像素级分割任务中,构建高-低质量的掩码偏好对,引导方法在隐空间学习精准的视觉表征.利用SAM(Segment Anything Model)的交互特性构建两类负样本:为了解决边缘不精细问题,在真值包围盒内引入对抗性点提示,生成局部缺失或溢出的低质量掩码作为边缘负例;为了解决目标定位错误问题,在背景区域随机采样生成非重叠掩码,构建语义级定位负例.经过多样本偏好对的训练,结合SAM实现高精度的掩码分割.在多个数据集上实验表明,FG-DPA可有效抑制定位幻觉,显著提升掩码生成的完整性与边缘准确度,在提升多模态生成式视觉感知性能方面是有效的.

【Abstract】 Generative referring segmentation methods based on multimodal large language model(MLLM) are limited by the mechanism of Supervised Fine-Tuning and lack in-depth exploration of ways to improve generation quality.Therefore,these methods are faced with the challenges of semantic localization bias and rough mask boundaries in complex scenarios.To address these issues,a fine-grained direct preference alignment framework for generative visual perception(FG-DPA) is proposed.The direct preference optimization (DPO) algorithm is transferred from text understanding to the pixel-level segmentation task.High-quality and low-quality mask preference pairs are constructed to guide the method toward learning more accurate visual representations within the latent space.Two types of negative samples are produced by leveraging the interactive characteristics of the segment anything model(SAM).To address the issue of imprecise edges,adversarial point prompts are introduced into the ground-truth bounding box to generate low-quality masks with local omissions or overflows as negative examples.To solve the problem of incorrect target localization,non-overlapping masks are randomly sampled in the background region to construct semantic-level negative examples.Through training with multiple samples,accurate segmentation is finally achieved in conjunction with SAM.Experiments on multiple public datasets show that FG-DPA effectively suppresses localization hallucination and significantly improves the completeness and edge accuracy of mask generation,validating its effectiveness in enhancing multimodal generative visual perception performance.

【基金】 国家自然科学基金项目(No.62073101);中国博士后科学基金项目(No.GZC20252736);机器人技术与系统国家重点实验室(哈尔滨工业大学)自主课题项目(No.SKLRS202417B,SKLRS202501C);安徽省机器视觉检测重点实验室开放基金项目(No.KLMVI-2024-HIT-18);黑龙江省“揭榜挂帅”科技攻关项目(No.2023ZXJ01A02)资助~~
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2026年03期
  • 【分类号】TP391.41;TP18
  • 【下载频次】1
节点文献中: 

本文链接的文献网络图示:

本文的引文网络