节点文献

粗细粒度因果关系协同驱动的可解释性视觉问答方法

Fine-to-Coarse Grained Causality Co-Driven Approach for Explanatory Visual Question Answering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 施业成缪佳李俞奎

【Author】 SHI Yecheng;MIAO Jiali;YU Kui;School of Computer Science and Information Engineering,Hefei University of Technology;Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education of China,Hefei University of Technology;

【通讯作者】 俞奎;

【机构】 合肥工业大学计算机与信息学院合肥工业大学大数据知识工程教育部重点实验室

【摘要】 可解释性视觉问答(Explanatory Visual Question Answering, EVQA)在回答视觉问题的同时为推理过程生成用户友好的多模态解释,从而提高模型推理的可信度.然而,由于缺乏对视觉区域对象关系的有效建模,现有EVQA生成的解释文本存在视觉区域与语义不一致的问题.为此,文中提出粗细粒度因果关系协同驱动的可解释性视觉问答方法(Fine-to-Coarse Grained Causality Co-Driven Approach for Explanatory Visual Question Answering, FCGC-CoD).首先,建模视觉区域特征的因果关系,识别其中的主体对象和支撑对象,增强视觉与语言预训练模型的多模态表征能力.然后,设计联合变分推理网络,通过细粒度的多模态因果表征增强模型粗粒度的宏观因果推理过程,实现多模态解释和答案的生成.实验表明,FCGC-CoD在准确回答问题的同时,可提升解释的视觉推理一致性.

【Abstract】 Explanatory visual question answering(EVQA) generates user-friendly multimodal explanations for the reasoning process while answering visual questions. Thereby, the credibility of model inference is enhanced. However, due to the lack of effective modeling of visual regions object relations, the explanations generated by existing explanatory visual question answering(EVQA) models suffer from the problem of inconsistency between visual regions and semantics. To address this issue, a fine-to-coarse grained causality co-driven(FCGC-CoD) approach for explanatory visual question answering is proposed. First, the causal relationships of visual regions features are modeled, and the influential and supportive objects are identified to enhance the multimodal representation capability of the vision-and-language pretrained model. Then, a joint variational causal inference network is designed to strengthen the coarse-grained reasoning process through fine-grained multimodal causal representations, and thus the generation of multimodal explanations and answers is achieved. Experimental results demonstrate that FCGC-CoD enhances the visual reasoning consistency of explanations while answering questions accurately.

【基金】 新一代人工智能国家科技重大专项项目(No.2021ZD0111801);国家自然科学基金项目(No.62376087)资助~~
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2025年06期
  • 【分类号】TP391.41
  • 【下载频次】17
节点文献中: 

本文链接的文献网络图示:

本文的引文网络