节点文献

基于信息瓶颈理论和时空特征融合的多摄像头协同推理研究

Research on Multi-camera Collaborative Inference Based on Information Bottleneck Theory and Spatiotemporal Feature Fusion

【作者】 张志国;

【导师】 张俊星;

【作者基本信息】 内蒙古大学 , 计算机科学与技术, 2025, 硕士

【摘要】 随着边缘计算与人工智能技术的深度融合,多摄像头协同视频分析在智能交通、智慧安防等领域展现出广阔的应用前景。然而,传统的云端集中式推理架构在处理高清视频时,往往面临通信带宽消耗大、端到端延迟高及系统可扩展性差等诸多瓶颈,难以满足实时性与低功耗的实际需求。尽管边缘计算通过将部分计算下沉至靠近数据源的终端设备,在降低延迟和保护隐私方面具有天然优势,但受限于边缘设备的计算与存储资源,其在复杂视频分析任务中的性能仍难以令人满意。针对上述挑战,本文提出了一种基于信息瓶颈理论与时空特征融合的多摄像头协同推理系统。该系统通过任务导向的数据压缩与编码,最大化地保留对视频分析任务最具价值的信息,从而显著降低通信成本,提高视频分析效率。同时,设计了一种轻量化的特征提取网络,结合高效多尺度注意力与结构优化策略,确保在资源受限的边缘设备上也能高效地提取关键视觉特征。为进一步提升多摄像头间的协同推理能力,本文提出了视图贡献加权(VCW)机制和Swin-T时空信息融合网络,分别在空间与时间维度对特征融合过程进行动态加权与全局时空建模,从而显著增强系统对复杂动态场景的适应性和鲁棒性。在Wildtrack和Multiview X两个公开多视角数据集上的实验表明,本研究在保持高检测精度的前提下,显著降低通信开销,端到端分析帧率达到了边缘部署的实际需求。尤其在MODA(Multiple Object Detection Accuracy,多目标检测准确率)、Precision(精度)、推理延迟和通信带宽占用等关键指标上,优于现有的基线方法。总的来说,我们提出的多摄像头协同推理方案在几个关键方面都有好的表现:它在保持较高检测精度的同时,推理速度快(延迟低),使用的网络资源也少(网络开销低),而且能在不同环境下稳定工作(环境适应性好)。这套方案为将来智能边缘场景下的多角度监控和高效视频分析打下了实用的基础。

【Abstract】 With the deep integration of edge computing and artificial intelligence technologies,multi-camera collaborative video analysis has shown broad application prospects in intelligent transportation,smart security,and other fields.However,traditional cloud-based centralized inference architectures face several bottlenecks when processing high-definition video,including high communication bandwidth consumption,high end-to-end latency,and poor system scalability,making it difficult to meet real-time and low-power requirements.Although edge computing offers inherent advantages in reducing latency and protecting privacy by offloading computation to devices close to the data source,its performance in complex video analysis tasks is still unsatisfactory due to the limited computational and storage resources of edge devices.To address these challenges,this paper proposes a multi-camera collaborative inference system based on information bottleneck theory and spatiotemporal feature fusion.The system uses task-oriented data compression and encoding to maximize the retention of the most valuable information for video analysis tasks,significantly reducing communication costs and improving video analysis efficiency.Additionally,a lightweight feature extraction network is designed,combined with efficient multi-scale attention and structural optimization strategies,ensuring the efficient extraction of key visual features on resource-constrained edge devices.To further enhance the collaborative inference capability between multi-cameras,this paper introduces a View Contribution Weighting(VCW)mechanism and a Swin-T spatiotemporal fusion network,which perform dynamic weighting and global spatiotemporal modeling of feature fusion in the spatial and temporal dimensions,respectively,significantly improving the system’s adaptability and robustness to complex dynamic scenarios.Experiments on two public multi-view datasets,Wildtrack and Multiview X,demonstrate that our method significantly reduces communication overhead while maintaining high detection accuracy,with end-to-end analysis frame rates meeting the practical requirements for edge deployment.Notably,it outperforms existing baseline methods in key metrics such as MODA(Multiple Object Detection Accuracy),Precision,inference latency,and communication bandwidth utilization.In conclusion,the multi-camera collaborative inference system proposed in this paper achieves innovative breakthroughs in several aspects:it not only maintains high detection accuracy but also achieves low inference latency and low network overhead,exhibiting excellent environmental adaptability.This system provides a solid technical foundation for multi-view perception and efficient video analysis in future intelligent edge scenarios.

  • 【网络出版投稿人】 内蒙古大学
  • 【网络出版年期】2026年 03期
  • 【分类号】TP391.41
节点文献中: 

本文链接的文献网络图示:

本文的引文网络