节点文献

面向交通场景的轨迹-图像跨模态检索方法研究

Research on Cross-Modal Trajectory-Image Retrieval Methods for Traffic Scenes

【作者】 张旭东;

【导师】 张蕊;

【作者基本信息】 武汉理工大学 , 软件工程, 2023, 硕士

【摘要】 交通是现代城市生活的重要一环,交通轨迹和路况图像作为交通数据的两种重要形式,它们之间的相互检索在城市交通管理、智能交通等领域的运用场景很有前景。基于深度学习的多模态检索方法具有检索精准度高、存储占用小和检索快速等优点。据本文所知,该研究尚处于初始阶段,且缺乏相关的数据集。而且,交通场景演变速度快、辨识难度高且轨迹位置点数据在时间上连续、在空间上相邻,捕捉其内在特征表示具有难度。对此,本文研究如下:首先,由于该领域的研究匮乏,缺少相关的交通多模态数据集,因此本文构建了一个包含轨迹、图像以及作为标签的描述文本的数据集。数据组成为成都滴滴出租车一个月的轨迹数据、交通路况图像以及相应的描述语句。最后从描述语句与图像、轨迹的相关程度对数据集进行评估,评估从四个方面进行:语法正确性、描述准确性、逻辑连贯性和语言流畅性。经过相关的专业人员评估,验证了数据集的有效性,可用于本文提出的任务。其次,由于城市交通场景的多模态数据具有很强的时序性,一般的深度学习网络对轨迹和图像的时空特性的挖掘能力不足,且大多基于全局的多模态语义对齐,数据的关联性与相似性难以捕捉,基于循环神经网络的模型对时序数据的特征提取存在信息压缩和无法解决信息长期依赖等问题。为此,本文提出了基于多模态注意力机制的交通数据检索方法,利用注意力机制的重点信息捕获能力,多个模态数据使用注意机制计算相似度系数,从而实现多模态数据细粒度对齐与后续的数据融合。在MAP和Recall@k指标上的表现说明了该模型相比其他模型有更好的表现。最后,针对城市交通场景多模态数据特征捕获细粒度的不足。本文提出了基于多模态Transformer的多监督检索方法。并且为了进一步提高模型的检索精度,本文的模型利用标签信息构造n×n的实例-实例语义相似度矩阵,指导模型的哈希码学习过程,其在实例-实例层面上弥补了异构数据带来的语义鸿沟问题。将多个模态表示输入Transformer,利用其他模型的信息增强单一模态的信息表示,使特征更具有代表性。实验结果显示,该方法在MAP和Recall@k指标上优于对比方法,拥有更好的检索性能。

【Abstract】 Traffic is an important component of modern urban life.Traffic trajectory and road condition images are two important forms of traffic data.The mutual retrieval between them has great potential for application in urban traffic management,intelligent transportation,and other fields.Deep learning-based multimodal retrieval methods have advantages such as high retrieval accuracy,small storage footprint,and fast retrieval.As far as this thesis knows,this research is still in its initial stage,and there is a lack of related datasets.Moreover,the traffic scene changes rapidly,and the identification difficulty is high.The trajectory location point data is continuous in time and adjacent in space,making it difficult to capture its intrinsic feature representation.This thesis proposes the following research:Firstly,due to the lack of research in this field and the absence of relevant multimodal traffic datasets,this thesis constructs a dataset that includes trajectories,images,and descriptive text as labels.The data comes from one month of trajectory data of Chengdu taxis,traffic road condition images,and corresponding descriptive sentences.Finally,the relevance between the descriptive sentences and images/trajectories is evaluated from four aspects: grammatical correctness,descriptive accuracy,logical coherence,and language fluency.After evaluation by relevant professionals,the effectiveness of the dataset was verified,and it can be used for the proposed task in this thesis.Secondly,due to the strong temporal characteristics of multimodal data in urban traffic scenes,conventional deep learning networks have insufficient ability to explore the spatiotemporal properties of trajectories and images.Also,they are based on global multimodal semantic alignment,and the association and similarity of data are difficult to capture.Models based on recurrent neural networks have problems in feature extraction for temporal data such as information compression and inability to solve long-term dependencies.Therefore,this thesis proposes a traffic data retrieval method based on multimodal attention mechanism.The focus information capture ability of attention mechanism is used to calculate the similarity coefficients of multiple modal data with the attention mechanism,thereby achieving fine-grained alignment of multimodal data and subsequent data fusion.The performance on MAP and Recall@k indicators shows that this model performs better than other models.Finally,to address the deficiency of fine-grained feature capture for multimodal data in urban traffic scenes,this thesis proposes a multimodal supervised retrieval method based on multimodal Transformer.In order to further improve the model’s retrieval accuracy,this model uses label information to construct an instance-instance semantic similarity matrix to guide the model’s hash code learning process,which bridges the semantic gap between heterogeneous data at the instance-instance level.The representations of multiple modalities are input into the Transformer,and the information from other models is used to enhance the representation of a single modality,making the features more representative.The experimental results show that this method performs better than the comparison methods on MAP and Recall@k indicators and has better retrieval performance.

  • 【分类号】U495;TP391.41
节点文献中: