节点文献

水声系统中基于强化学习的子空间信道跟踪

Reinforcement Learning Based Channels Tracking by Tracking the Signal Subspace in Underwater Acoustic Systems

【作者】 王宇航;

【导师】 李维;

【作者基本信息】 哈尔滨工业大学 , 信息与通信工程, 2022, 硕士

【摘要】 近年来海洋通信需求日益增加,建立高效稳定水下通信系统的意义随之显现。但海洋环境复杂多变,导致水声信道具有快速时变特性并存在严重的时频域双拓展。准确可靠的信道跟踪可以纠正由水声信道快速时变特性产生的信号幅度衰落和相位偏移,使系统的误码率满足实际应用需求,有效可靠实现水声通信。信道的快速时变特性为信道跟踪过程带来了诸多困难,强化学习凭借其自主学习优势在应对这一难题上有很大潜力,通过强化学习应对时变特性带来的影响已成为水声信道跟踪领域中的趋势。研究表明,根据水声信道的交叉路径相关性,可以将信道跟踪转变为对低维子空间和相应主成分的跟踪。然而这种跟踪算法中存在两个问题。首先,算法的设计没有考虑快速时变特性对相关性的影响,将子空间维度看作时不变参数,由此造成的模型失配问题会大大影响跟踪算法的性能;其次,在跟踪信道主成分的过程中,受快速时变特性影响,信道模型无法准确描述环境特征,导致卡尔曼滤波器的精度降低,收敛性变差。本文在国内外研究成果的基础上,针对上述两个问题展开研究。针对时不变子空间维度所导致的模型失配问题,提出了基于强化学习的时变维度子空间信道跟踪。考虑到信道的相关性受时变特性影响,强化学习智能体根据环境的变化实时调整子空间维度。将时变维度的强化学习模型,与传统的子空间信道跟踪算法相结合,通过设置合适的子空间维度使信道模型匹配信道环境,以应对快速时变特性造成的信道模型失配问题。针对卡尔曼滤波对信道模型失配的容忍度较差问题,提出了基于强化学习的双向自适应子空间信道跟踪。自适应双向卡尔曼滤波采用所有时刻的数据对跟踪结果进行平滑,同时通过遗忘因子约束卡尔曼增益,调整输入数据对跟踪结果的校正程度,降低了突发数据的影响。采用强化学习算法调整遗忘因子门限值,通过设置动态的门限值,应对快速时变特性的影响,提升跟踪算法的准确性和稳定性。本文通过使用AUVFest 07海试的真实实验数据,证明了无论在平静海况还是恶劣海况下,所提出算法凭借强化学习的自主学习优势,在快速时变信道中都能保持较高的准确性与稳定性,使系统误码率平均降低约20%。

【Abstract】 To meet the growing demand of underwater acoustic communication(UAC),it is important to establish an efficient and stable underwater acoustic(UWA)system.However,the ocean environment is complex and volatile,resulting in UWA channels with fast time-varying characteristics and severe expansion in the time-frequency domain.The effective and reliable implementation of UAC requires accurate and efficient channel tracking to correct the random phase shift and amplitude fading generated during signal transmission,so that the bit error rate(BER)of the system can meet the practical requirements.The fast time-varying nature of the channel poses many difficulties for the channel tracking process,and reinforcement learning(RL)has great potential to address this challenge with its autonomous learning advantage.It has become a trend in the field of UWA channel tracking to cope with the effects of fast time-varying characteristics through RL.It has been shown that channel tracking can be transformed into tracking of lowdimensional channel subspace and corresponding principal components based on the cross-path correlation of UWA channels.However,there are two problems in this tracking algorithm.First,the design of the algorithm does not consider the influence of fast time-varying characteristics on the correlation between the channel taps,setting the subspace dimension as a time-invariant parameter,and the resulting model mismatch problem will greatly affect the performance of the tracking algorithm.Second,in the process of tracking the channel principal components,the channel model accuracy decreases due to the influence of fast time-varying characteristics,resulting in the reduction of the accuracy of the Kalman filter and the convergence deterioration.Based on the recent research results,this work addresses the above two problems.An RL-based time-varying dimension subspace channel tracking is proposed to address the model mismatch problem caused by the time-invariant subspace dimension.Considering that the correlation of channels is affected by time-varying characteristics,the RL intelligences adjust the subspace dimension in real time according to the changes of the environment.The time-varying dimensional RL model is combined with the traditional subspace channel tracking algorithm to cope with the channel model mismatch problem caused by the fast time-varying characteristics.An RL-based forgetting factor forward-backward filtering subspace channel tracking is proposed to address the problem of poor tolerance of Kalman filtering to channel model mismatch.The forgetting factor forward-backward Kalman filtering uses data from all time slots to smooth the tracking results,while constraining the Kalman gain to adjust the degree of correction of the input data on the tracking results and reduce the influence of burst data.The RL algorithm is used to adjust the forgetting factor threshold value to improve the accuracy and stability of the tracking algorithm by setting a dynamic threshold value to attenuate the effect of time-varying characteristics.By using practical experimental data from the AUVFest 07 sea trial,this work demonstrates that the proposed algorithms can maintain high accuracy and stability in fast time-varying channels with the autonomous learning advantage of RL,reducing the system BER by about 20% on average,no matter in calm or rough sea conditions.

  • 【分类号】TN929.3;TP181
节点文献中: