节点文献
新型多元校正、校正转换和多元分类分析方法研究
Novel Approaches to Mutivariate Calibration and Classification and Calibration Transfer
【作者】 倪网东;
【导师】 满瑞林;
【作者基本信息】 中南大学 , 化学工艺, 2010, 博士
【摘要】 现代化工过程和化工产品质量涉及多方面因素和指标,通过这些因素和指标的测定数据来关联化工过程和化工产品质量,对优化整个化工过程和产品质量具有重要意义。由于系统往往非常复杂,这些因素和指标往往具有相关性,测定数据混有噪音或干扰。因此从大量复杂数据中提取有用信息、滤除噪音,建立稳健的数学模型,可获得对于生产过程更好的理解,实现对产品质量进行在线控制。本文选择广泛用于质量在线控制过程的近红外光谱系统为研究对象,研究并提出了几种化学计量学方法,涉及信号处理、噪音滤除、稳健模型建立、仪器漂移校正等方面。本论文首先考察两种纯净分析信号或处理(Net Analyte Signal/Processing, NAS/NAP)算法、两种正交信号校正(Orthogonal Signal Correction, OSC)算法,研究了它们之间及其与PLS多元校正之间的关系。实验数据和模拟光谱数据处理结果表明,在一般情况下,这些信号处理方法处理光谱后得到的多元校正模型的预测结果相似;但在光谱中的噪声和不同组分之间存在重叠的条件下,根据其结果可将这些方法分成两组:应用Lorber NAP算法建模与应用FearnOSC算法建模获得相同的预测结果;而应用Olivieri NAP算法建模与应用修改后的Fearn OSC算法建模获得相同的预测结果。这四种信号处理方法仅能简化PLS模型,而不能提高PLS模型的预测效果。因此,本论文开发了新型分段NAP信号处理方法--Piecewise NAP (PNAP),通过局部地去除光谱中与目标分析物不相关的干扰或噪声,从而提高在近红外光谱上建立的多元校正模型的预测效果。与常用的OSC或NAP信号处理方法相比,即使光谱中包含噪声或不同组份之间存在重叠,PNAP算法模型的预测结果比常规NAP和OSC模型更稳健。与Olivieri的NAP和改良Fearn的OSC一样,对Olivieri的NAP和改良Fearn的OSC进行分段处理(Piecewise),其多元校正模型具有相同的预测结果。基于模型融合思想,本文开发了叠加偏最小二乘回归(Stacked Partial Least-Squares, SPLS)和叠加移动窗口偏最小二乘回归(Stacked Moving-Window Partial Least-Squares, SMWPLS)两种新型多元校正方法,它们通过给予大小不同的权重因子来去除光谱中的局部冗余信息。其基本原理是在所有的光谱间隔上建立平行的常规偏最小二乘回归(PLS)模型,充分利用整个光谱数据中的信息,并且通过赋予光谱数据中与目标分析物高度相关光谱间隔较大的权重,对这些子PLS模型进行融合。理论和实验结果表明这两种叠加算法效果显著优于传统的PLS模型。这两种叠加算法可以获得更简化的回归模型。当目标分析物的信息不是均匀地分布或仅集中在一个光谱间隔上时,与在整个光谱或一个最好的光谱区间内建立的PLS模型相比,叠加模型可以获得更好的预测效果。另外,应用这两种叠加算法不仅可以提高多元校正模型的预测效果,而且对于光谱中包含离群点(Outlier)的数据具有潜在的处理能力。本文研究同样显示新开发的叠加融合方法(SPLS)具有保持多元校正模型的能力,可将某分析仪器建立的校正模型用于另一分析仪器的分析预测上。与传统的校正转移方法需要在两个分析仪器上都进行转换标准样本检测不同,SPLS算法不需要测量转换标准样本,就可在两个不同仪器的光谱数据上都可以获得稳健预测效果,并保持良好的模型预测能力。通过湖中沉积物的近红外光谱数据,展示SPLS对不同仪器上样本的预测效果。SPLS通过减小不同仪器上的局部光谱差异的影响,从而有助于常用的校正转换技术提高转换后模型的预测效果。当在频率空间中进行校正转换时,基于数据或模型融合思想的双域回归分析(Dual Domain Regression Analysis, DDRA)方法同样有助于提高常规的校正转换模型的预测效果。在不同的空间中(时域和频域)进行模型融合都可以提高校正转换后的模型的预测效果,而本文新开发的SPLS与常用的校正转换方法相结合优于利用双域回归(DDRA)模型融合的校正转换方法。与回归中一样,对分类器进行融合或叠加,去除光谱中局部的冗余信号,相对于任何一个单独的分类器,可以获得更精确的分类结果。本文开发了两种新的叠加分类器法,包含一个基于新设计SPLS的叠加偏最小二乘分类分析方法(Stacked Partial Least-Squares Discriminant Analysis, SPLSDA)和另一个对一系列线性分类器进行叠加的叠加线性分类分析方法(Stacked Linear Discriminant Analysis, SLDA)。与偏最小二乘分类分析(Partial Least-Squares Discriminant Analysis, PLSDA)和线性分类分析(Linear Discriminant Analysis,LDA)分类器的分类结果相比,应用SPLSDA和SLDA可获得更好的分类结果。SPLSDA分类器通过对在不同的光谱间隔上建立的子PLSDA分类器进行叠加融合,从而揭示光谱中的每个光谱间隔对最终分类结果的贡献,并且SPLSDA一般比在整个光谱上建立的PLSDA分类器需要更少的潜在变量,SPLSDA和SLDA分类器使用不同的权重进行变量选择,同时弥补了在常规的变量选择方法中可能发生的信息遗漏缺陷。本论文还研究了利用小波正交信号校正(Wavelet Orthogonal Signal Correction, WOSC)去除频率空间中的不相关信息,从而获得稳健的分类器。这个新的分类方法将Wavelet Prism(具有局部的和多频率组分的特点)与正交信号校正(整体地去除不相关的信息)相结合,从而极大地提高分类器的分类效果,同时简化了分类器。本文的研究显示小波正交信号处理分类分析(Wavelet Orthogonal Signal Correction Discriminant Analysis, WOSCDA)分类器可以有效地去除光谱中与分类不相关的信息,与在整个时域上进行OSC处理的PLSDA分类器(Orthogonal Partial Least-Squares Discriminant Analysis, OPLSDA)相比,WOSCDA分类器的效果更好。
【Abstract】 Many contemporary chemical manufacturing processes and quality of their products are involved in many factors and indices, which are significantly important to the optimization of chemical manufacturing processes and quality of their products. It is theorial importance and high valuable to apply that extracting real information and filtering noise from complex data to build robust mathematical models, which is benifical for better understanding of manufacturing process and better on-line quality control of final product, because of complexity of manufacturing processes, high colinerarity of different factors or indices and existence of interference or noise in obtained data.Near infrared spectroscopy (NIRS) with wide application in on-line quality and process control was selected in the thesis as research object to propose several kinds of multivariate calibration methods, including signal processing, noise filtering, robust model building and baseline correction.In this dissertation, firstly, a comparison was made between two versions of Net Analyte Signal/Processing (NAS/NAP) algorithms and two different Orthogonal Signal Correction (OSC) algorithms to reveal the relationship between these four methods and that between them and PLS model. Although the NAP and OSC algorithms preprocessed spectra differently, we showed by comparison of the external prediction errors (RMSEP) of these algorithms based on real and synthetic data that these two types of algorithms had the same predictive performance. Under some conditions, the overlap and noise, slightly differentiated these algorithms:the performance of Lorber’s NAP algorithm and Fearn’s OSC algorithm tracked closely and that Olivieri’s NAP algorithm tracked with the performance of a modified form of Fearn’s OSC. All these four signal processing methods couldn’t improve predictive performance from PLS model, but simplify it. Therefore, a novel signal processing method, Piecewise Net Analyte Processing (PNAP), was developed to enhance the predictive performance from multivariate calibration model based on NIR spectrum through local removal of unrelated information. And through comparison of predictive performance, PNAP was obviously superior to traditional OSC and NAP, even though the spectra contained noise and overlapping of different components. Like Olivieri’s NAP and modified version of Fearn’s OSC, it was also shown piecewise implementations of Olivieri’s NAP and the modified version of Fearn’s OSC to filter a set of spectra had the same predictive performance.Two novel algorithms which employed the idea of stacked generalization or stacked regression, Stacked Partial Least-Squares (SPLS) and Stacked Moving-Window Partial Least-Squares (SMWPLS) was reported to remove local redundant information through unevenly distributed weights. The new algorithms established parallel, conventional PLS models based on all intervals of a set of spectra to take advantage of the information from the whole spectrum by incorporating parallel models in a way to emphasize intervals highly related to the target property. It was theoretically and experimentally illustrated that the predictive ability of these two stacked methods was never poorer than that of a PLS model based only on the best interval. These two stacking algorithms generated more parsimonious regression models with better predictive power than conventional PLS, and performed best when the spectral information is neither isolated to a single, small region, nor spread uniformly over the response. Additionally, the work about these two stacked methods did not only demonstrate the improvement, but also demonstrate that stacked regressions had the potential capability of predicting property information from an outlier spectrum in the prediction set.In this dissertation, we also showed the capability of stacked methods to maintain the predictive performance from calibration model. Unlike transfer methods requiring measurement of transfer standard on the primary and secondary instruments, SPLS regression can be used to generate parsimonious regression models with good predictive power on both primary and secondary instruments, without calibration transfer in most cases. The predictive performance from SPLS for predicting samples between different instruments was demonstrated through lake sediment NIR data set. Conventional calibration transfer techniques could also take advantage of stacked PLS regression to minimize local instrumental differences between two instruments. Dual Domain Regression Analysis (DDRA) based on the same idea of data fusion might also contribute to the improvement in predictive performance from conventional calibration model, when calibration transfer methods were implemented in frequency domain. Data fusion in different domains could enhance the predictive performance from transferred model, but fusion using SPLS was much better.Classifiers in combination and fused classifiers removing redundant information locally, like that in calibration, might generate more accurate classification than any single classifier. In this dissertation, we developed a few new stacked classifiers, including Stacked Partial Least-Squares Discriminant Analysis (SPLSDA), a classifier based on SPLS and two approaches to Stacked Linear Discriminant Analysis (SLDA), a classifier that combines stacking with linear discriminant analysis. It was shown in this dissertation that improvement in classification performance obtained after application of stacked PLSDA and stacked LDA as compared with that obtained using PLSDA and LDA classifiers. A stacked PLSDA classifier developed by weighting and combining local classifiers built on separate regions of the data revealed the contributions of those regions of the data to the classification, and often required fewer latent variables for same classifications than a conventional PLSDA classifier applied to the whole data set. The stacking weights generated in stacked PLSDA and stacked LDA performed a kind of variable selection while compensating for information loss that might occur in conventional variable selection techniques.In this dissertation, we used Wavelet Orthogonal Signal Correction (WOSC) for multivariate classification through removal unrelated information in frequency domain. This new classification tool combined a wavelet prism decomposition of a spectral response (local and multi-scale property in frequency scale) and orthogonal signal correction (global filtering uncorrelated classification information) to significantly improve the classification performance either in term of reduction of classification errors and in reduction of model complexity. We showed that a discriminant analysis based on WOSC removed irrelevant classification information effectively and performed favorably as compared to a wavelength-domain filtering approach, such as that used in Orthogonal Partial Least-Squares Discriminant Analysis (OPLS-DA).