节点文献

视频运动人体行为识别与分类方法研究

Human Motion Recognition and Classification in Video Sequences

【作者】 张浩

【导师】 刘志镜;

【作者基本信息】 西安电子科技大学 , 计算机应用技术, 2011, 博士

【摘要】 运动人体行为的特征表示和目标识别作为视频监控的主要研究内容,是计算机视觉领域当前的研究热点,不但具有重要的实际意义,而且对计算机视觉和其它相关研究领域具有重要的促进作用。视频监控技术研究的主要目的是将人类的视觉感知功能赋予机器视觉系统,使其能够在图像序列中发现目标、跟踪目标,并对目标的行为进行识别和理解。经过几十年的不懈研究,上述技术均取得了长足的进步,但实践表明通常意义上的人体行为检测与识别技术还远未成熟,有效的人体行为识别算法是智能视频监控系统鲁棒性和实用性的核心。研究工作以视频监控系统为载体,以实际应用为研究目标,以运动人体行为作为主要研究对象,以行为的辨别方法为研究内容,深入地探索了人体行为识别与分类方法中行为的表示和建模方法。研究内容涵盖了单人与多人行为、局部与全局特征、基于模板的方法和基于概率模型的方法,为智能视频监控系统的实用化提供理论依据。本文的主要工作和贡献概述如下:(1)分段加权动态时间规划法研究了步态模型中多模板之间的关联性,针对多模板之间关联程度弱、不易度量的问题,提出了一种新的多模板相似性度量方法,即分段加权动态时间规划法。在分析步行运动规律的基础上,总结了步行过程中身体形状的宽高比与各个姿态的对应关系。按照姿态的相似程度,将一个完整的步态周期划分为八个连续的状态,从每个状态中分别选取连续的三帧图像作为特征帧。然后提取特征帧的步态轮廓,计算轮廓的质心点坐标,利用解卷绕法将二维轮廓线展开为一维信号,统计该信号量的长度,分析其统计特征,确定标准化的参数值。将所有解卷绕后的数据标准化并插值,获得每个数据点的值,并计算相同状态一维信号的算术平均值,按顺序构建对应状态的模板。计算其各状态模板与样本的加权动态时间规划距离,度量相似程度,判别测试序列所属类别,提高了步态识别的准确率。在解卷绕过程中,为了精确地计算质心到轮廓边界点之间的距离,提出了单像素轮廓点提取算法,解决了由于轮廓线像素宽度不统一而引起的展开误差问题。该方法在建模过程中,充分分析了数据规律,并做了必要的数据规范和插值,克服了因为数据尺度差异产生的误差。(2)关键帧特征识别法研究了运动人体行为姿态的差异性,针对单人运动行为过程复杂、难于表征的问题,提出了一种新的行为模板识别方法,即关键帧特征识别法。通过对人体运动行为规律的研究,分析了各种行为的特征姿态,总结了不同行为身体形状的宽高比与对应特征姿态的关系,提出了一种特征姿态的检测算法。该算法消除了噪声点和关键帧误差的影响,有利于从视频序列中提取有效信息,减少冗余信息,降低计算复杂度,提高计算效率。通过对?变换描述符的研究,总结了其在图像描述方面的性质,利用该描述符分析了视频序列中各种姿态的特征曲线及变化规律,通过计算同类特征姿态的算术平均值建立了关键姿态模板,降低了误差。该方法充分利用关键姿态差异大的特点,采用动态时间规划方法计算模板与测试样本之间的距离,度量相似程度,判别测试样本所属行为类别。它克服了序列数据之间刚性比较的缺点,实现了柔性比较,提高了行为识别的准确率。(3)关键点特征分类法研究了运动人体行为特征的简化表示问题,针对单人运动行为特征数据冗余、计算复杂度高的问题,提出了一种新的行为特征简化表示与分类方法,即关键点特征分类法。通过研究人体运动行为的规律,分析了行为轮廓表示法的缺点,总结了运动行为的特点,提出了关键点特征表示法,减少了噪声的干扰,克服了轮廓特征数据冗余和维数高的缺点,达到了行为特征简化表示的目的。分析了隐马尔可夫模型的结构和参数关系,根据特征向量的维数关系,表示了行为的初始迭代参数,提出了观测概率的表示方法,解决了特征向量维度过高的问题,降低了计算复杂度。利用Baum-Welch算法迭代估计参数,通过Viterbi算法估计产生最大概率的路径,利用内积度量距离,解决了因样本数量有限而产生的参数估计欠精确问题,提高了人体行为分类的准确率。(4)交互行为整体分类法研究了两人交互行为的整体表示问题,针对交互行为中两人相互关系不易表示的问题,提出了一种新的两人交互行为分类方法,即交互行为整体分类法。通过研究两人交互行为的过程,总结了交互行为的变化规律,根据交互过程中两人之间的联系紧密程度划分了三个阶段,分析了各阶段行为的特点,提出了交互行为特征整体表示法,避免了在个人行为表示的基础上分析两人的相互关系,简化了交互行为特征的表示过程,降低了误差。通过分析隐马尔可夫模型的结构和参数依赖关系,设置了行为模型的初始参数,利用Baum-Welch算法迭代估计参数,通过Viterbi算法估计产生最大概率的路径,达到判别交互行为类别的目的。该方法有效地保持了时变数据的顺序关系,解决了模板法中时变数据间关联性弱的问题,提高了交互行为分类的准确率。(5)整体特征判别分类法研究了运动人体行为各姿态之间的关联性问题,针对隐马尔可夫模型各状态间独立性弱的问题,提出了一种新的运动人体行为分类方法,即整体特征判别分类法。研究了人体运动行为的变化规律,总结了各种行为的姿态特征,结合?变换描述符在图像方面的性质,分析了各种行为姿态的?变换曲线及变化规律,归纳了曲线与行为变化规律之间的联系。然后研究了线性链式条件随机场的结构,比较了其与条件随机场的差异,利用了其线性链式结构关系,以及隐状态之间的相对独立性,实现了分类多种单人运动行为。该方法不但有效地利用了时变数据的顺序关系,而且解决了隐马尔可夫模型中时变数据之间独立性弱的问题,提高了行为分类的准确率。以上五种方法不但在人体行为数据库上得到了验证,而且在实际应用过程中获得了较好的效果。

【Abstract】 Feature representation and object recognition are centric for video surveillance, and hot topics in computer vision. They effectively promote computer vision and other related fields, and are practical research questions in real world applciaitons. The final goal of video surveillance is to transplant human vision perceptiion to machine, by enabling the video surveillance system to detect and track object, recognize/understand its behavior in video sequences. Though computer vision technologies have improved remarkably in recent years, detecting and recognizing human activities are still not matured. Therefore, intelligent algorithms in human activity recognition need to be studied to develop robustness and practice intelligent video surveillance systems.This research work explores activity representation and modeling in human recognition and classification. We utilize video surveillance as a carrier to study human activities recognitions. The content of this thesis includes single and multiple human activities, local and global features, template-based and probability-based approachs. Our work presents new approaches for intelligent video surveillance in real world application.The main contributions in the dissertation are described as follows.(1) Individually Weighted Dynamic Time WarpingBased on studies of relationship among mutliple templates in gait model, we present a novel approach named individually weighted dynamic time warping, which can measure the similiarity of multiple templates and solve the problem that different templates are less dependent and comparative. Firstly, we calculate ratios of width to height in body shape and summarize the relationship between them and each posture after have analyzed walking rule. Secondly, a total gait cycle is divided eight continual states by similiar postures, and three adjoining frames in each state are selected as feature frames. Thirdly, we extract gait silhouette from these frames, calculate the position of its mass center, unfold the silhouette to one-dimensionality signal, add the signal value, analyze their statistics and finally obtain normalized scale. Fourthly, the unfolded data are normalized and interpolated, and new values are obtained. Then we compute averages of one-dimensionality signal in the same state and construct eight state templates sequentially to be a gallery. Finally, we also obtain probes’templates by this method, calculate and compare the dynamic time warping distance between the probe and each gallery, and finally improve correct recognition accuracy. To compute the distance between mass center and contour, we propose an algorithm of extracting exact pixel in the contour in unfolding and solve the problem that errors result from the irregular width of contour. This method sufficiently employs the rules, necessarily normalizes and interpolates the data, and avoids the errors of irregular data in size.(2) Keyfame-based feature recognitionBased on studies of diverse postures in human motion, we propose a new method named keyframe-based feature recognition, which is template-based activity recognition and solves the problem that motion processes are diverse and hard to be described. Firstly, we analyze key-posture in diverse activities and summarize the relationship between them and each posture after have studied the rules in human motions. This approach avoids noises and keyframe errors, extracts features from video sequences conveniently, eliminates redundant information, reduces computational complexity and finally improves computational efficiency. Secondly, after ? transform descriptor is discussed, its properties are analyzed in image representation, and then it is employed to analyze feature curves and rules in video sequences. Thirdly, we compute arithmetical average of feature postures in the same categorty to construct key-posture template so that the errors are reduced. Finally, this approach focuses on diverse key-postures, and employs dynamic time warping to calculate the distances between each galleries and probe to measure similarity, so that it recognizes motion categorty. It implements flexible distances and avoids computing Euclidean distance between two sequences, so that correct recognition accuracy is improved.(3) Keypoint-based feature classificationBased on studies of compact representation in human motion, we present a novel approach named keypoint-based feature classification, which is compact representation and classification in activity features and solves the problem that it is redundant in motion features and high in computational complexity. Firstly, we analyze shortages in silhouette-based representation and summarize motion characteristics after having studied the rules in human motions. This method not only avoids noises and redundant data, but also achieves the goal of compact representation. Secondly, we discuss the structure and parameters in Hidden Markov Model, initialize iterative parameters by the dimensitionality of feature vectors, and propose observation representation to solve the problem in high dimensitionality and computational complexity. Finally, the parameters are estimated iteratively by Baum-Welch algorithm. The Viterbi algorithm is employed to estimate the path in the highest probability. The inner product is utilized for distance measurement. This method solves the problem that parameter estimation is inexact because of lack in instances, so that correct classification accuracy of human motion is improved.(4) Interaction classification of global featuresBased on studies of global representation of two subjects’interactions, we propose a new method named interaction classification of global features, which solves the problem that the relationship between two subjects is prone not to be described. Firstly, we summarize the rules after having studied two subjects’interaction, divide three phases to analyze their characteristics, and propose global representation for interactions to avoid analyze their relationship after single subject representation. Thus it simplifies the stages in representation and reduces errors. Secondly, the structure and parameters in Hidden Markov Model are discussed and initializations are decided. Finally, we employ Baum-Welch algorithm to estimate them iteratively and Viterbi algorithm to obtain the path in the highest probability, so that we classify the interactions. This method preserves the relationship in sequential data sufficiently, solves the problem of independency in template-based approaches, and finally improves correct classification accuracy in interactions.(5) Discriminative classification of global featuresBased on studies of the relationship among human postures, we present a novel classification approach in human motions named discriminative classification of global features, which solves the problem that adjoining states are less independent. Firstly, we summarize posture features in each motion after having studied the rules in motions, discuss properties of ? transform descriptor in image representation, analyze its curves in diverse postures and their rules, and conclude their relationship. Secondly, Linear Chain Conditional Random Field is discussed, and the differences are analyzed in comparison with classic Conditional Random Field. Finally, we adopt its characteristics in structure and independency among hidden states to classify human motions. This method not only employs the relationship in sequential data, but also solves the problem of less independency in Hidden Markov Model, and finally improves correct classification accuracy in human motion.

  • 【分类号】TP391.41
  • 【被引频次】14
  • 【下载频次】1844
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络