节点文献
一般分布模式下GIS位置数据的不确定性研究
The Research on the Uncertainty of Position Data in GIS in General Distribution Mode
【作者】 游扬声;
【作者基本信息】 武汉大学 , 大地测量学与测量工程, 2005, 博士
【摘要】 GIS的不确定性已成为当前国际GIS界公认的重点研究课题。空间点元、线元的位置不确定性是GIS位置不确定性理论研究的基本内容。长期以来,人们在GIS不确定性研究领域进行了不懈的探索,取得了丰硕的成果,积累了大量经验。人们已经注意到,GIS位置数据的误差有可能不是正态分布的,而目前GIS位置数据不确定性研究一般是在误差分布是正态的基础上进行的。GIS位置数据的不确定性研究如果能顾及误差的实际分布模式,可以得到更好的成果。 本文的研究围绕GIS位置数据的误差分布可能是非正态的这个问题展开。不假设误差分布是正态的,会引出系列问题,例如:如何对GIS位置数据误差的分布做出合理的假设;如何估计误差的实际分布;如何描述误差的实际分布与正态分布的差异或者接近程度;如何度量非正态分布的数据的质量;如何建立非正态分布的数据的误差传播模型等等。本文对上述问题进行了较为深入和系统的研究。本文的主要内容如下: 第二章研究数字化数据的误差分布及相关问题。第1节率先推导了点元的坐标误差相关的情况下的位置误差的分布密度,目的是回答GIS用户常常会提出的点位误差平均有多大的这类问题。为了减弱或消除由于图纸变形、扫描数字化变形引起的GIS空间数据的系统误差,需要对数字化地图进行纠正。第2节讨论了数字化地图纠正的概念模型,相似变换、仿射变换、多项式回归模型均是描述该概念模型的数学模型;研究了纠正模型的检验和优选,并用实例考察了纠正模型的优选及纠正的效果:第3节提出了用计算机伪随机数生成p—范分布样本的快速方法,并基于p—范分布样本,研究了误差分布的拟合检验方法,也由此引出误差分布估计的研究。 针对误差分布的估计,第三章以信息扩散估计(核估计)的理论研究为中心。第1节提出了扩散估计的最优窗宽的概念和实用算法;第2节提出了全局最优窗宽的概念、理论和估计方法,基本上彻底解决了扩散估计的关键的如何确定窗宽的问题,丰富了扩散估计的理论,改善了扩散估计的效果,为扩散估计成为研究误差分布的经典方法奠定了坚实的理论基础。第3节提出了基于扩散估计的误差分布的解析拟合方法,优于基于直方图估计或者经验分布的误差分布的拟合检验方法:第4节提出了基于信息扩散的极大似然估计理论和方法,该方法具有根据实际误差分布自动选择估计方法的能力,丰富了研究误差分布、参数估计的方法库;在单参数极大似然估计中,给出了一个不依赖于显著水平的比较客观的判断粗差的标准。 第四章提出了误差分布的相似度的理论,对不同类型的问题,给出了相似度的计算方法。本章的成果为分布的估计、拟合提供了合理的、自然的准则;为定量研究误差分布的相似、近似、渐近等概念以及由于误差分布的近似引起的数据处理误差奠定了理论基础。 第五章研究GIS位置数据的不确定性度量。针对目前GIS位置数据一般没有协方差阵信息的实际情况,第2节提出了在协因数或权阵未知的情况下利用GIS叠置中的同名点元估计方差的方法。第3节提出了叠置层方差分量的极大似然估计方法,顾及了统计量的相关性,顾及了统计量的分布——维希特分布,采用极大似然法估计了各叠置层(同名点元)的方差分量,在位置误差服从正态分布的条件下,改进了第2节的研究成果。第2节、第3节的贡献在于不需要已知叠置层间的协因数阵或权阵。评价误差分布非正态或未知的数据的质量,熵是一个较合理的指标。第4节研究了误差分布的一般模式下的熵不确定度。指出了连续随机变量熵的数值计算问题,提出了基于扩散估计的熵不确定度的实用算法,导出了计
【Abstract】 The uncertainty of GIS is now one of the academic focuses in the community of GIS all over the world. The uncertainty of the position of spatial point and line segment underlies the theoretical research in the uncertainty of the position of GIS. For a long time, persistent efforts have been devoted to the research in the uncertainty in GIS. Great achievements and abundant experience have been harvested satisfactorily. It is noticed that the errors of the GIS position data are probably not normally distributed while most of the academic research in the uncertainty of GIS position data is being conducted on the basis of normal distribution of errors. Better achievements can be expected if the actual distribution mode of errors is taken into consideration in the research of uncertainty in GIS position data.This paper centers on the assumption that the actual distribution of the errors in GIS position data might be non-normal. A series of problem will arise without the assumption that the distribution of the errors is normal, such as: how to reach a reasonable assumption on the distribution of the errors of GIS position data? How to estimate the actual distribution of the errors? How to describe the differences or similarities between the actual distribution of the errors and the normal distribution? How to measure the quality of the non-normally distributed data? How to model the error propagation for non-normally distributed data? This paper presents a thorough systematic study of the preceding problems. The distribution of the errors of digitized data and the related issues are researched in the second chapter . In the first section the author deduces the distribution density of position errors in relation to coordinate error of points in the hope that it can serve as an answer to the frequent question from GIS users how the errors of points vary onthe average. In order to diminish and eliminate the systematic error of spatial data inGIS caused by the deformation of the map and digitized scanning, correction should be done to a digitized map. The second section discusses the conceptual model of thecorrection of a digitized map. Similarity transformation, affine transformation and polynomial regression model are all mathematical models describing the conceptual model. This section focuses on test and optimization of the correction models and a case study is conducted to test the optimization of a correction model and the effect of the correction. The third section presents an easier way of generating P-normdistributed sample with the help of pseudo-random number form computers. Based on the P-norm distributed samples, academic research is conducted on the methods of collocation testing of error distribution, with further research of the estimation of error distribution.Oriented to further study of Error Distribution Estimation, Chapter Three dwells on theories of Information Diffusion Estimation. The first section presents the concept and algorithm of Optimal Window-width of Diffusion Estimation. The second section puts forward the concept, theory and estimation techniques of overall optimal window-width, which enriches the theories of Diffusion Estimation with much more reliable estimating results as an essential solution to ascertaining the window-width and lays a solid theoretical foundation in making Diffusion Estimation the best-preferred technique in the research of error distribution.The third section presents the Analytic Collocation of Error Distribution based on Diffusion Estimation. Guided by this technique, we can get the unique analytic expression of an error distribution, which works better than the Collocation Testing methods of distribution based on the histogram or the cumulative distribution. The fourth section presents the theory and techniques of Maximum-Likelihood Estimation based on Information Distribution. This technique is able to automatically choose the best estimation method according to the actual error distribution, thus consequently enriches the variety of techniques in the study of error distribution and parameter estimation. In Maximum-Likelihood Estimation of single parameter, an objective standard for judging gross error is suggested independent of a prominent level.Chapter Four advances the similarity theory of error distribution and presents a calculating method of similarity concerning different kinds of problems. The achievements in this chapter offer a logical and natural guide rule for the estimation and analytical collocation of distribution, and serves as a theoretical foundation for the quantitative study of similarity, approximation and graduality of error distribution as well as the study of error of data-processing caused by the approximation of error distribution.In face of the fact that there is currently no covariance matrix information for GIS position data, the second section in Chapter Five offers a new technique of estimating variance by using corresponding points in overlaying of GIS when the weight matrix is unknown. The third section advances Maximum-Likelihood Estimation of the overlay layers’ variance components, considering the correlation ofstatistic and its distribution (Wishart Distribution). In this way, the Maximum-Likelihood is adopted to estimating the variance components of each map layer (corresponding points), and ameliorates the achievements mentioned in section two when positional errors are normally distributed. The striking contribution of the second and third sections is that we don’t have to know the correlation matrix or weight matrix of the overlay layers when we want to successfully evaluate the quality of the non-normal distributing or unknown distributing data. Entropy is a more appropriate index, especially when the variance of the data’s error distribution does not exist. The fourth section offers a research on the uncertainty of entropy of error distribution in general distribution mode. The author points out the issue how to calculate the numerical value of the entropy of a continuous random variable, propounds the practical arithmetic of entropy-uncertainties on the basis of Diffusion Estimation, and educes the relation between calculating entropy’s integral interval and interceptive error limit. When error is normally distributed, error ellipse is satisfactory enough to serve us as the standard tool to portray the randomicity of points on a plane. The fifth section presents the concept and theory of error oblate with which the author depicts the randomicity when positional error of points on planes obeys the P-norm distribution. It is the natural extension of error ellipse, and it is of definite statistic significance.The sixth chapter presents a series of study of the propagation of GIS position data error in the mode of general distribution. Currently, there is generally no distribution mode attached to GIS position data. By taking advantages of the contrast of overlaying and the analysis of statistics in the hope of finding out the statistical distribution of the position data, this chapter dwells on error distribution of point coordinates in the analysis of overlaying. The second section offers a deduction of the difference distribution between the positions of corresponding points and offers a detailed study of analytical collocation of error distribution when the similarity of distribution reaches the maximum. Since the difference between the positions of corresponding points is true error, the research based on it is of great significance. The third section studies the estimation of the error distribution of the positions of corresponding points. With the knowledge of general mode of error distribution, by restricting the solution to unknown function is a density function of P-norm distribution, the estimation of the distribution of the corresponding points’ positional error is offered and the shortcut approximation solution based on the criterion ofMaximum-Similarity is advanced. Corresponding points represent the spatial-characterized points, thus the distribution from estimation supplies information for reference to the positional error distribution of other characteristic points. P-norm-Maximum-Likelihood Estimation is introduced to estimate the corresponding points’ coordinates in the overlaying export map in the Section Four.The error band model of line segments in GIS is the difficult part of significant importance in the study of the uncertainty of position data in GIS. Most existing error bands are constructed based on points, that is to say, constructed in a certain way on the basis of error distribution of a random point on a line segment. The error band model based on points can hardly show us the probability of the line segments’ true position falling within the error band. The seventh chapter puts forward the concept and theories of integral equi-density error band of a line segment in GIS, constructs the integral euqi-density error band of line segments, provides the short-cut approximation arithmetic concerning the probability of the line segments’ true position falling within the error band, and offers here some relevant diagrams. Furthermore, based on the theories of error oblate, this dissertation constructs an integral equi-density error band model in the mode of general error distribution and lays a foundation for the practical application of error band of a line segment in GIS.
【Key words】 Error Distribution; Information Diffusion Estimation; Similarity; Error Oblate; Integral Equi-Density Error Band;