节点文献

基于文本数据驱动与语言模型增强的多模态理解方法研究

Text-Driven and Language Model-Enhanced Multimodal Understanding

【作者】 李威;

【导师】 杨易;

【作者基本信息】 浙江大学 , 计算机科学与技术, 2025, 博士

【摘要】 随着人工智能技术的不断进步,多模态学习作为实现计算机视觉与自然语言处理深度协作的核心技术,已成为人工智能领域的前沿研究方向和关键挑战之一。尽管近年来多模态领域取得了显著发展,但现有模型在复杂多模态上下文理解与深层次推理等方面仍面临突出困难,极大地限制了其在现实场景中的广泛应用。针对上述挑战,本文提出了一种基于“文本数据驱动”与“语言模型增强”的新型多模态学习路线。具体而言,文本作为人类对世界抽象描述的天然载体,内在地包含了丰富的视觉概念与语义信息,为多模态融合提供了天然的切入点。与此同时,大语言模型凭借其强大的通用知识、语义理解与推理能力,为多模态任务提供了高效且通用的认知引擎。本文将两者有机结合,克服了传统多模态学习依赖大量人工标注多模态数据的问题,使模型能够在更广泛、更复杂的多模态任务中实现高效泛化与深层推理。具体来说,针对现有多模态模型泛化能力差、严重依赖人工标注数据的问题,本文提出一种基于纯文本驱动的零样本视觉描述方法(De Cap)。该方法创新地利用CLIP模型高效对齐视觉与文本特征空间,通过纯文本数据训练图像描述模型,大幅降低了对人工标注数据的依赖,并显著提高了模型的零样本泛化性能。其次,针对现有模型在复杂图文交互理解任务中多模态推理能力不足的问题,本文进一步提出了多模态组合学习框架(MCL)。该框架通过引入多模态组合类型的预训练任务,显著强化了模型对复杂图文信息的深层理解与跨模态推理能力,从而有效提升了视觉问答、组合图文检索视觉故事生成等任务的性能。此外,针对视频理解任务中高质量标注数据稀缺与时序推理难度高的问题,本文进一步提出了一种基于纯文本预对齐的视频理解方法(TOPA)。TOPA通过自动生成的文本序列有效地模拟视频内容的动态变化,显著减少了对真实视频数据的依赖,并利用纯文本数据将视频模态与语言模型联结,实现了对视频内容更精细的表达与推理。总的来说,本文围绕“文本数据驱动”与“语言模型增强”这一创新路线,对从静态图像到动态视频、从基础理解到复杂推理的一系列多模态学习难题展开了研究。本文所提出的方法在多个实际任务中表现出卓越的泛化能力和显著的应用潜力,为未来多模态学习的发展提供了清晰有效的方法指引和丰富的实践依据。

【Abstract】 With the rapid advancements in artificial intelligence,multimodal learning,which inte-grates computer vision and natural language processing,has emerged as a frontier research area and key challenge in AI.Although significant progress has been made recently,existing models still struggle with complex multimodal contextual understanding and deep reasoning,severely limiting their applicability in real-world scenarios.Thus,enhancing multimodal mod-els’contextual comprehension and reasoning capabilities remains a central research question in multimodal learning.To tackle these challenges,this thesis proposes a novel multimodal learning paradigm driven by text data and large language models.Text inherently serves as a natural carrier of abstract human descriptions,containing rich visual concepts and semantic information,and thus provides an intuitive entry point for multimodal fusion.Simultaneously,large language mod-els(LLMs),with their powerful general knowledge,semantic understanding,and reasoning capabilities,offer a more efficient and versatile cognitive engine for cross-modal tasks.By effectively integrating these two components,the proposed paradigm overcomes the conven-tional dependency on large-scale annotated multimodal datasets,allowing models to generalize efficiently and reason deeply in diverse and complex multimodal contexts.Specifically,to address poor generalization and the heavy reliance on annotated data in existing vision-language models,this thesis first introduces a text-only zero-shot visual cap-tioning approach(De Cap).De Cap innovatively leverages the CLIP model to align visual and textual feature spaces effectively,enabling an image captioning decoder to be trained on text-only data,substantially reducing dependence on human annotations.Moreover,a training-free cross-modal feature projection mechanism is developed,mitigating modality discrepancies and significantly enhancing zero-shot generalization performance.Secondly,recognizing the limitations in existing models’cross-modal reasoning capabil-ities for complex image-text interactions,the thesis proposes a multimodal composition learn-ing framework(MCL).The MCL framework,through jointly designed tasks of multimodal-contextual captioning(MC-Cap)and multimodal-contextual retrieval(MC-Ret),considerably strengthens large language models’capabilities for deep comprehension and cross-modal rea-soning.This approach effectively improves performance across various multimodal tasks,in-cluding visual question answering,composed image retrieval,and visual storytelling.Additionally,addressing the scarcity of high-quality annotated data and challenges in tem-poral reasoning for video understanding,this thesis presents a text-only pre-alignment method(TOPA)for video understanding.TOPA innovatively constructs a Tideo representation by auto-matically generating textual sequences to simulate dynamic video content,significantly reduc-ing reliance on real video data.Through text-based pre-alignment,TOPA effectively connects video modalities with language models,harnessing LLMs’semantic comprehension and con-textual modeling abilities.In conclusion,guided by the innovative integration of text data and large language models,this thesis systematically develops a unified multimodal learning framework.It progressively addresses critical multimodal challenges,from static images to dynamic videos,and from fun-damental understanding to sophisticated reasoning.The proposed methods exhibit outstanding generalization capabilities and significant practical potential,providing clear methodological guidance and a robust foundation for future research in multimodal learning.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2026年 07期
  • 【分类号】TP391.1;TP391.41
节点文献中: