节点文献

汽车座舱智能交互关键技术研究及应用

Research and Application of Key Technologies for Intelligent Interaction in Automotive Cockpits

【作者】 张强;

【导师】 石琴;

【作者基本信息】 合肥工业大学 , 机械工程(专业学位), 2025, 博士

【摘要】 随着通信、计算机及人工智能的快速发展,汽车正从传统交通工具向智能移动终端转变。在此进程中,汽车座舱作为人机交互最直观的载体,已成为汽车产业技术升级的重点。在当前产业实践中,基于语音识别、机器视觉和手势感知的智能交互体系已初步构建,其便捷性、智能性有效降低了驾驶员的操控负荷,提升了乘客的驾乘满意度。然而,在车载环境下仍存在若干关键技术瓶颈。比如:在手势识别方面,受限于光照条件的复杂性和人员姿态的可变性,现有神经网络模型的准确率和鲁棒性仍显不足;在语义理解方面,非指令性语音存在较高的触发概率。这些问题本质上源于复杂环境下多源异构数据的特征提取能力不足,以及车载算力约束下的模型轻量化难题。基于作者在企业技术中心多年的工程实践经验,本文聚焦于多模态交互技术在汽车座舱中的创新应用,重点探讨轻量化约束下的手势识别、语音识别,驾驶意图理解及拒识模型等。主要如下:1)面向智能座舱的轻量化手势识别技术研究针对汽车座舱车载场景中视角受限、遮挡干扰及算力约束下的手势识别难题,本文提出了一种基于Nano Det的轻量化手势识别模型。该模型构建了“人手检测-关键点识别-手势识别”的时空联合感知框架,并应用了多维特征提取机制。在空间维度,采用无锚框(Anchor-free)检测策略,结合自适应训练样本选择(ATSS)和广义焦点损失(GFL),有效解决了类别不平衡和难易样本问题,使检测精度与鲁棒性显著提升;通过特征图和热力图拼接技术实现手部2D/3D关键点检测,在DO、ED、RHD和Assembly任务上的AUC分别达到95%、90.06%、88.8%和93.4%。在时序维度,通过计算帧内与帧间残差,捕捉手部关键点与动作信息,并结合全连接层进行特征融合与手势分类。在模型轻量化方面,采用三阶段优化策略。首先,通过深度可分离卷积替代传统卷积层,压缩了62.3%的参数量,其次,通过构建特征金字塔网络(FPN)和路径聚合网络(PAN)的级联结构,提取多尺度特征;最后,采用Res Net18替代Res Net50,进行关键点检测特征提取,在保持检测精度的同时,降低了模型复杂度和模型体积,提升了GPU推理速度(30.90ms),实现了实时性目标。测试结果表明,该模型在自主构建的包含9类标准手势的舱内测试数据集上表现优异,静态与动态手势识别准确率分别达到94%和96.5%。该研究为车载环境下的手势识别提供了高效、实时的解决方案,具有一定的工程应用价值。2)面向智能座舱的多模态语音识别技术研究针对车载环境的噪声干扰问题,本文提出了一种基于改进型AV Hu BERT框架的多模态语音识别模型。在轻量化的特征提取方面,采用可分离卷积(DSC)重构Res Net18特征提取网络。在跨模态的时空对齐方面,通过Transformer架构,建立唇部动作与语音信号的时空对齐机制。数据处理方面,构建了包含公开数据集与实车座舱环境数据的混合训练集。通过MFCC特征提取与K-means聚类(K=2000)生成伪标签,并采用基于唇音同步验证的两步筛选机制确保数据质量。领域迁移实验表明,混合数据预训练策略可以显著提升中文座舱场景的识别性能。该研究实现了车载环境下高精度、低延迟的语音识别,为智能座舱交互系统设计提供了一种解决思路。3)面向智能座舱的驾驶意图理解与多模态拒识技术研究针对拒识机制在车载环境中面临的挑战,如:方言及非指令性语音等,本文构建了车载场景专属的多模态交互数据集,并提出了基于大语言模型(LLM)的多模态拒识模型。在数据集方面,本文首次建立了融合“语言-行为-环境”多维特征的车机交互拒识数据集,数据采集涵盖实车测试中的语音记录、驾驶员面部朝向、手势动作及情绪状态等多模态信息。数据集构建采用四阶段标准化流程(采集-筛选-标注-划分),精准区分"人机指令"与"人人交流"两类核心标签,最终形成包含训练集与验证集的高质量数据集。在模型设计方面,本文以开源Chat GLM2-6B大语言模型为基础架构,结合"P-Tuning v2"参数高效微调技术,通过任务导向的提示(Prompt)设计实现通用语言知识向车载场景的领域适配。实验设计了生成式与抽取式两类提示模板,并对比了单模态(仅语音)与多模态(语音+非语言信息)输入对模型性能的影响。测试结果表明,多模态输入显著提升了模型的决策效能。面部朝向、车内人数等非语言信息为模型提供了关键上下文线索,有效增强了车载场景的驾驶意图判别能力。该研究通过构建首个车载场景多模态交互数据集,提出基于LLM的多模态拒识模型,为解决车载环境下的意图判别难题提供了新的解决方案,对提升智能座舱交互安全与用户体验具有一定参考价值。4)面向智能座舱的多模态人机交互测试验证平台开发针对智能交互技术在车载环境下的功能安全验证与性能量化评估难题,本文基于高通SA8255P芯片架构,构建了多模态人机交互测试验证平台。该平台集成了视觉感知模块(驾驶员手势/唇形识别、面部朝向分析、情绪检测)与语音交互模块(语音指令识别、自然语言处理),实现了本文多种模型框架在真实车载场景下的测试验证。系统测试包括三个典型场景:基于单目摄像头的车载手势识别(支持空调调节、音乐切换等6类功能)、基于声音-图像的车载语音指令识别(覆盖通信、导航、娱乐等10类场景),以及基于Chat GLM2预训练大模型的多模态驾驶意图判别(结合面部朝向、情绪状态、车内人数等特征)。系统硬件以高通SA8255P芯片为核心座舱域控制器,它支持多路音视频输入的,集成了DMS/OMS摄像头模组、麦克风阵列及中控显示系统。系统软件采用模块化设计,包含音频处理引擎(实现语音采集、降噪与ASR)、视觉计算单元(部署轻量化手势识别模型)及多模态融合决策模块(基于预训练大模型的意图理解)。测试结果表明,多模融合策略可以有效提升车载场景的人机交互体验,增强系统的鲁棒性和泛化能力。综上所述,针对汽车座舱智能交互技术在车载场景的应用瓶颈,本文通过“模型创新-数据构建-测试验证”的闭环研究思路,构建了轻量化的手势识别、语音识别网络模型,提出了基于Chat GLM2预训练大模型的拒识模型框架,搭建了多模态人机交互测试验证平台,开展了基于汽车座舱多模态交互专属数据集的测试验证,模型性能在实时性、准确性及场景适应性方面均有显著提升。相关研究成果可为汽车座舱智能交互系统的工程化落地提供一定的技术参考,并为未来高阶智能座舱的交互设计提供思路。

【Abstract】 With the breakthrough development of Internet of Vehicles,mobile computing,and artificial intelligence technologies,modern automobiles are undergoing a paradigm shift from traditional transportation tools to intelligent mobile terminals.In this process,the automotive cockpit,as the intuitive carrier of intelligence and the core field of human-machine interaction,has emerged as a strategic focal point for technological upgrading in the automotive industry.In current industrial practices,a fusion interaction system based on speech recognition,machine vision,and gesture perception has been preliminarily established.Through naturalized human-machine collaboration,this system significantly reduces driver cognitive load and enhances passenger satisfaction.However,under the human-machine co-driving paradigm,several critical technical bottlenecks remain:(1)Dynamic Gesture Recognition:Constrained by complex cockpit lighting conditions and variable driver postures,existing convolutional neural network models face significant trade-offs between real-time performance and accuracy.(2)Complex Semantic Understanding:Non-instructional speech semantic filtering mechanisms exhibit high false triggering probabilities.(3)Driving Intention Comprehension:Traditional finite state machine-based decision models struggle to effectively handle spatio-temporal asynchronous issues among multi-modal inputs,resulting in insufficient intention prediction accuracy.These challenges fundamentally stem from inadequate representation learning of multi-source heterogeneous data in complex scenarios and the dilemma of model lightweighting under vehicle system computing constraints.Drawing on the author’s years of engineering practice experience in the corporate technology center,this paper focuses on the innovative applications of multi-modal interaction technology in automotive cockpits,and mainly explores its technical paths in enhancing driving safety,optimizing user experience and reconstructing the relationship between humans and vehicles.The details are as follows:1)Research on Lightweight Gesture Recognition Model for Intelligent CockpitIn view of the challenges of gesture interaction in intelligent cockpit scenarios,such as limited perspectives,occlusion interference,and computing power constraints,this paper proposes a lightweight gesture recognition model based on Nano Det.The model constructs a spatio-temporal joint perception framework of"human hand detection-key point recognition-gesture recognition",and innovatively integrates a multi-dimensional feature extraction mechanism.In the spatial dimension,an Anchor-free detection strategy is adopted,combined with Adaptive Training Sample Selection(ATSS)and Generalized Focal Loss(GFL),which effectively addresses the problems of class imbalance and difficult/easy samples,significantly improving detection accuracy and robustness.Hand 2D/3D key point detection is achieved through the feature map and heatmap stitching technology,with AUC reaching 95%,90.06%,88.8%,and 93.4%respectively in the DO,ED,RHD,and Assembly tasks.In the temporal dimension,hand shape and motion information are captured by calculating intra-frame/inter-frame residuals,and feature fusion and gesture classification are carried out in combination with fully connected layers.In terms of model lightweight design,a three-stage optimization strategy is adopted to break through the bottleneck of computational efficiency.Firstly,depth-separable convolution is introduced to replace the traditional convolution layer,achieving a 62.3%parameter compression rate.Secondly,by constructing a concatenated structure of Feature Pyramid Network(FPN)and Path Aggregation Network(PAN),multi-scale features are extracted.Together with the lightweight detection head,it completes object classification and bounding box regression,ensuring multi-scale hand detection results.Finally,the key point detection module uses Res Net18 to replace Res Net50 for feature extraction.While maintaining high detection accuracy,it reduces the model’s computational complexity and size,optimizes the GPU inference speed to 30.90ms,and achieves real-time performance.Real vehicle test results show that the model excels on the self-built cabin test dataset containing eight standard gestures(such as thumbs up,V-sign,heart-shaped gesture,upward wave,downward wave,left wave,right wave,and grab).The accuracy of static and dynamic gesture recognition reaches 94%and 96.5%respectively.In complex lighting and multi-pose scenarios,the recognition accuracy remains above 90%,meeting the reliability requirements of automotive standards.This research provides an efficient and real-time solution for natural interaction in the vehicle environment,which has certain engineering application value.2)Research on Multimodal Speech Recognition Technology for Intelligent CockpitAddressing the issues of noise interference and computational resource limitations in the vehicle environment,this paper proposes a multimodal speech recognition model based on the improved AV Hu BERT framework.This model constructs a visual-auditory dual-modal fusion architecture.In terms of lightweight feature extraction,the depth-separable convolution(DSC)is used to reconstruct the Res Net18video feature extraction network.In cross-modal alignment,a spatial-temporal alignment mechanism between lip movements and speech signals is established by introducing the Transformer architecture.In data processing,a mixed training set(200 hours in duration)was constructed,including the LRS3 public dataset and real-car cockpit environmental data.MFCC feature extraction and K-means clustering(K=2000)were used to generate pseudo-labels,and a two-step screening mechanism based on lip-sync verification was adopted to ensure data quality.Domain migration experiments showed that the mixed data pre-training strategy could significantly improve the recognition performance in Chinese cockpit scenarios,with the word error rate(WER)decreasing from 100.5%of the pure English baseline model to 33.33%.In the NVIDIA Tesla V100 GPU environment,the model’s video feature extraction parameters were reduced to 1/9 of the original architecture,with a 20.3%reduction in fine-tuning time and a 6.58%improvement in WER.This research achieved high-precision and low-latency speech recognition in complex vehicle environments,providing a solution for the design of intelligent cockpit interactive systems.3)Research on Driving Intent Understanding and Multi-modal Recognition Technology for Intelligent CockpitAddressing the significant challenges faced by pure language-based rejection mechanisms in traditional in-vehicle speech interaction systems,particularly in complex environments with noise interference,non-standard language expressions,and the diversity of drivers’non-linguistic behaviors leading to high false recognition rates and limited interactive experiences.This paper innovatively constructs a multi-modal interactive dataset specific to in-vehicle scenarios and proposes a multi-modal rejection model based on large language models(LLMs).In terms of datasets,this paper establishes an in-vehicle interaction rejection dataset that integrates"language-behavior-environment"multi-dimensional features for the first time.Data collection covers multi-modal information such as voice recordings during actual vehicle testing,driver facial orientation,gesture movements,and emotional states.The dataset construction follows a four-stage standardization process(collection,screening,annotation,and partitioning),ensuring privacy security through encrypted storage and double manual verification.The annotation process innovatively combines speech semantics and non-linguistic behavioral features to accurately distinguish between two core labels:"human-machine instructions"and"human-human communication,"ultimately forming a high-quality dataset containing training and validation sets.In terms of model design,this paper uses the open-source Chat GLM2-6B large language model as the basic architecture and combines efficient fine-tuning technology with the"P-Tuning v2"parameters to achieve domain adaptation from general language knowledge to in-vehicle scenarios through task-oriented prompt design.The experiment designs both generative and extractive prompt templates and compares the impact of single-modal(voice only)and multi-modal(voice+non-linguistic information)input on model performance.Test results show that multi-modal input significantly improves the decision-making effectiveness of the model.With the optimal prompt design,the multi-modal model achieves an accuracy rate(ACC)of 95.6%,an F1 score of 96.0%,and a false rejection rate(FRR)reduced to 2.8%,achieving a leapfrog improvement compared to the pure language model(ACC 89.5%,FRR 10.2%).Feature contribution analysis shows that non-linguistic information such as facial orientation and the number of people in the car provides critical contextual cues for the model,effectively enhancing intent discrimination in complex scenarios.This research provides a new solution to the problem of intent discrimination in in-vehicle environments by constructing the first multi-modal interactive dataset for in-vehicle scenarios and proposing a multi-modal rejection model based on LLMs.It has certain reference value for improving the interactive safety and user experience of smart cockpits.4)Development of a multi-modal human-machine interactive testing and validation platform for intelligent cockpitsAddressing the challenges of functional safety verification and performance quantification evaluation of multi-modal technology under complex working conditions,this paper designs and builds a multi-modal human-machine interaction testing and validation platform based on the Qualcomm SA8255P chip architecture.The platform integrates visual perception modules(driver gesture/lip recognition,facial orientation analysis,emotion detection)and voice interaction modules(voice command recognition,natural language processing).Based on the speech processing capability of the cloud big model and the edge-side computer vision algorithm,this platform realizes the testing and validation of multiple model frameworks in real car-borne scenarios,through the voice processing capabilities of cloud-based large models and edge-side computer vision algorithms.System testing covers three typical scenarios:in-vehicle gesture recognition based on a single camera(supporting six functions such as air conditioning adjustment and music switching),in-vehicle voice command recognition based on sound and image(covering ten scenarios such as communication,navigation,and entertainment),and multi-modal driving intention discrimination based on the pre-trained Chat GLM2 large model(combined with features such as facial orientation,emotional state,and number of people in the car).The system’s hardware architecture centers on the Qualcomm SA8255P chip,building a cockpit domain controller that supports multiple audio and video inputs,integrating DMS/OMS camera modules,microphone arrays,and central display systems.The software architecture adopts a modular design,including an audio processing engine(for voice collection,noise reduction,and ASR),a visual computing unit(deploying a lightweight gesture recognition model),and a multi-modal fusion decision module(based on the pre-trained large model for intention understanding).Test results show that through a multi-modal fusion strategy and training with a dedicated dataset,the multi-modal rejection model’s ability to distinguish instructions in complex environments has significantly improved.This demonstrates the application potential of multi-modal information fusion in improving in-vehicle interaction experience and verifies the crucial role of cross-modal information complementarity in enhancing the robustness of in-vehicle interaction.In summary,aiming at the application bottlenecks of intelligent interaction technology in automotive cockpits,this paper adopts a closed-loop research paradigm of"algorithm innovation-data-driven-system verification".It constructs a lightweight multi-modal fusion architecture,proposes a driving intention understanding framework based on the pre-trained large model Chat GLM2,builds a multi-modal human-machine interaction test and verification platform,and conducts tests and verifications based on the exclusive multi-modal interaction dataset for automotive cockpits.The model performance has been significantly improved in terms of real-time performance,accuracy,and scene adaptability.The relevant research results can provide certain technical references for the engineering implementation of intelligent interaction systems in automotive cockpits and offer ideas for the interaction design of future high-level intelligent cockpits.

  • 【分类号】U463.6
节点文献中: 

本文链接的文献网络图示:

本文的引文网络