节点文献
基于麦克风阵列的声源定位系统的研究与实现
Research and Implementation of Sound Source Localization System Based on Microphone Array
【作者】 黄云;
【导师】 王浟;
【作者基本信息】 电子科技大学 , 电子信息(专业学位), 2023, 硕士
【摘要】 近年来,得益于语音交互市场的快速发展,麦克风阵列声源定位技术逐渐成为研究热点。麦克风阵列声源定位技术是实现语音增强、滤除噪声的有力手段,并被广泛应用于智能家居、会议系统和工业检测噪音等各大领域。但当前的定位系统大多基于PC机完成数据处理和位置解算,体积大、集成度低,并且系统的应用环境常常存在噪声和混响,传统定位算法还有待改善。针对这些问题,本文提出了基于小型麦克风阵列和ZYNQ开发平台的声源定位方案,定位算法基于先进行时延值估计,后进行位置估计的TDOA双步定位法并做改进,本文所做的主要工作如下:(1)从理论上阐述语音信号的处理方法并对影响麦克风阵列性能的参数进行仿真分析。对语音分帧、加窗和活动性检测方法进行阐述,为后续使用算法处理语音信号提供理论基础。接着采用CBF算法对影响阵列性能的因素进行仿真分析,为阵列的设计提供思路。(2)针对传统时延估计算法在混响和低信噪比环境下性能较差的问题,提出了一种改进时延估计算法。首先建立IMAGE混响模型,对传统时延估计算法的常见加权函数进行仿真,并提出了基于PHAT加权和二次互相关算法相结合的改进时延估计算法。在IMAGE混响模型中的仿真结果表明,改进时延估计算法在混响和低信噪比环境下的时延估计均方根误差要低于传统算法。(3)针对TDOA算法在位置估计时时延误差容易被放大的问题,基于最小二乘法设计了一种位置估计算法。本文首先建立了多基线定位模型并确定了麦克风阵列的几何结构,接着推导出位置参数与时延估计值以及时延估计误差间的方程组,并基于最小二乘法设计迭代算法。最后对定位算法的均方误差与克拉美罗下界(Cramer-Rao Lower Bound,CRLB)进行推导,证明了角度估计精度接近CRLB。(4)基于ZYNQ开发平台完成对声源定位系统的设计与实现。首先对麦克风阵列及其接口进行硬件设计,接着在PL部分设计实现多路数据的处理和时延值估计,为了将时延值送入PS部分进行角度计算,对PL与PS部分的通信系统进行设计实现,最后在PS部分实现定位算法并输出定位结果。声源定位系统的测试结果表明,系统完成单次定位的平均耗时为23.47 ms,在3 m距离内的方位角与俯仰角平均绝对误差不超过3°,满足参数指标要求,并对比相关文献,验证了定位系统在定位实时性和精度上的优势。
【Abstract】 In recent years,sound source localization(SSL)technology based on microphone array(MA)has gradually become a research hotspot with the rapid development of the speech interaction market.SSL technology based on MA is a powerful means to achieve speech enhancement and noise removal,which is widely available for smart homes,conference systems,industrial noise detection sniper and other major fields.However,most of the current localization systems are based on PC platform to process data and realize localization,which has the disadvantages of large volume and low integration.And the traditional SSL algorithms need to be further improved because the application environment is often accompanied by noise and reverberation.In this thesis,a SSL system based on small MA and the ZYNQ platform is proposed to solve the problems raised above.The SSL algorithm is improved based on the TDOA,which is a two-step localization method of time delay estimation(TDE)before localization estimation.The main contributions in this thesis are as follows.(1)The processing method of speech data is introduced theoretically and the parameters affecting the performance of MA are simulated and analyzed.In order to facilitate the subsequent algorithm to process the speech signal,the speech framing,windowing and activity detection methods are expounded.The factors affecting the performance of MA are simulated and analyzed based on the CBF algorithm,which provides ideas for designing MA later.(2)Propose an improved TDE algorithm to solve the problem of poor performance of the traditional algorithm under the conditions of reverberation and low signal-to-noise ratio.The IMAGE reverberation model is established firstly,and the common weighting functions of the traditional algorithm are simulated.An improved TDE algorithm based on PHAT weighted and second cross correlation is proposed.The results of simulation in the IMAGE reverberation model prove that the root mean square error of the improved TDE algorithm under the conditions of reverberation and low signal-to-noise ratio is lower than that of the traditional TDE algorithm.(3)Design the localization algorithm according to the least squares method to solve the problem that the TDE error of TDOA algorithm in position estimation is easy to be amplified.This thesis first establishes a multi-baseline localization model and determines the MA geometry.After deriving the equation system between the localization parameter,the TDE value and error,the iterative algorithm is designed based on the least squares method.Finally,the mean square error and the Cramer-Rao Lower Bound(CRLB)of the localization algorithm are derived,and it is proved that the accuracy of angle estimation is close to CRLB.(4)Design and complete the SSL system based on the ZYNQ platform.Firstly,the hardware design of MA and its interface are carried out,then the multi-channel speech data processing and TDE are designed in the PL part.In order to input the time delay value into the PS part for angle calculation,the communication system between the PL and PS part is designed and implemented.Finally,the localization algorithm is implemented in the PS part and the angle of speech are output.The test results of the SSL system show that the system takes an average of 23.47 ms to complete a single localization,and the average absolute error of the azimuth and elevation angles do not exceed 3° within a range of 3 meters,which meets the requirements of the parameter indicators.The advantages of SSL system in real-time localization and accuracy are verified compared with the relevant literatures.
【Key words】 Microphone Array; Sound Source Localization; TDOA; Time Delay Estimation Algorithm; ZYNQ;
- 【网络出版投稿人】 电子科技大学 【网络出版年期】2024年 05期
- 【分类号】TN912.3