节点文献

人口统计数据空间化不同方案及其误差评价

The Different Schemes And Error Evaluation on Spatialization of Statistical Population Data

【作者】 杨旭

【导师】 廖顺宝;

【作者基本信息】 河南大学 , 地图学与地理信息系统, 2015, 硕士

【摘要】 人口数据统计一般是以行政区划为单元进行的,因此,在一个行政区域内部,人口只能被看成是均匀分布的(尽管实际情况并非如此),这就为利用人口数据进行跨学科的综合研究带来了不便。人口统计数据空间化技术有效地解决了这一问题。自人口数据空间化概念提出以来,国内外研究工作者已经在人口数据空间化方面做了大量的研究,提出了各种各样的方法和模型。其中,以人口分布与土地利用、土地覆被、夜间灯光等要素之间的统计方法和模型居多,这种现象在近十多年来的国内人口数据空间化研究中表现得特别明显。作为地学数据处理与产品加工的一种方法,与其他地学数据处理与产品加工方法一样,人口统计数据空间化也必然会存在误差。当前,关于人口数据空间化误差的分析与评价,研究人员一般都是针对自己特定的空间化方案和空间化结果进行,就事论事,但对于不同空间化方案(如选择不同参数、采用不同的统计样本尺度等)之间的误差分析、评价和比较所进行的系统分析和研究尚未发现。而系统分析和研究人口数据空间化不同方案之间的误差及其变化情况,对于优选空间化模型方法,减少空间化误差,提高空间化产品的精度,均具有重要意义。基于此,本文以2005年全国县级行政区划图、2005年分县统计人口数据以及全国1:25万土地覆被数据为基本资料,开展了以下研究:(1)空间化方案设计。(a)分析样本尺度的选择:根据人口分布与土地覆被类型之间的关系,从1:25万土地覆被数据的6个一级类型、25个二级类型中选择森林、草地、农田、城镇聚落、农村聚落、湿地水体、荒漠等7种土地覆被类型的面积作为回归分析的输入变量,人口统计数据为输出变量,基于县级和地市级统计样本尺度分别构建出1种人口数据空间化模型,两个尺度上共2种模型。(b)分析变量的选择:将与人口分布密切的农田、城镇聚落、农村聚落等3种土地覆被类型的面积作为固定参数,其他土地覆被类型(森林、草地、湿地水体、荒漠)的面积作为可选参数,在县级和地市级尺度上分别构建出14种模型,两个尺度上共28种模型。(2)误差评价指标与评价方法。建立了模型评价指标和空间化产品误差评价指标体系。本文选取相关系数、拟合值、相对误差、绝对误差等作为空间化模型和空间化产品误差评价的主要指标。同时,针对统计模型很难避免计算结果出现负值这一情况,除上述评价指标以外,对空间化结果的评价还考虑了计算结果中栅格单元值为负的占比情况。(3)误差分析。(a)以森林、草地、农田、城镇聚落、农村聚落、湿地水体、荒漠等7种土地覆被类型面积为参数,对基于县级统计样本和地市级统计样本的2种人口统计数据空间化方案进行比较,结果为基于县级单元样本的人口统计数据空间化方法较好,相关系数为0.797,相对误差为6.2%;(b)对县级、地市级两个尺度上28种不同参数的人口统计数据空间化模型进行分析、比较,结果为:以农田、城镇聚落、农村聚落、湿地水体的面积为参数,在县级尺度上的人口统计数据空间化方法较好,相关系数为0.797,相对误差为8.3%。(4)方案选择与优化。根据误差分析结果选择最佳方案,并对其进行优化。(a)基于县级尺度模型,在全国范围内分区进行人口统计数据空间化。依据相关系数、回归系数、散点图等指标,采用删除散点图上异常值的方法优化各区模型,使各个区的模型达到最好的结果,相关系数都在0.92以上,相对误差为0.134%。(b)对基于县级统计样本、以农田、城镇聚落、农村聚落、湿地水体面积为参数的人口统计数据空间化模型进行优化,参数的数量保持不变。依据相关系数、回归系数、散点图等指标,采用删除散点图上异常值的方法,使回归分析的结果达到最好,相关系数为0.955,相对误差为0.131%;(c)分析比较(a)、(b)两种方法的空间化结果,对其进行误差评价,最后得到:基于县级统计单元、以农田、城镇聚落、农村聚落、湿地水体面积为参数的人口统计数据空间化方法最优。本文通过不同方案的比较、分析,得到人口数据空间化的最优方案。通过该方案,使人口分布能更加直观地表现在空间上,更好地反映人口分布的实际情况,提高了人口数据空间化的精度和准确性。因此,本文的研究思路、研究方法和研究成果对今后的人口统计数据空间化具有一定的指导和参考作用。

【Abstract】 The statistical population data are generally recorded by administrative units. Therefore, inside an administrative zone, population can only be regarded as uniformly distributed though the actual situation is not so. Population data with this structure is not suitable to comprehensive analysis among different study areas. However, the spatialization technology of statistical population data effectively solves this problem. Since the concept of spatialization of population data was proposed, domestic and foreign researchers have done a lot of research in the fields of spatialization of population data, and put forward various kinds of methods or models. Most parts of these methods or models are based on the relationship between population distribution and land use, land cover, night lights or other related factors, this kind of situation looks particularly obvious for domestic researcher in past ten years.As a kind of method of geoscience data processing and products out putting, spatialization of statistical population data inevitably results in errors like other methods. So far, the analysis or evaluation of errors on spatialization of population data generally occurred at specific specialization models or specialization results that the researchers developed themselves. However, systematic researches, analysis, evaluation and comparison of errors among different spatialization schemes have not been found, for example, errors of spatialization models based on different parameters or scales of statistical samples. It is beneficial to preferring models of spatialization, reducing errors and improving precision of spatialization product to analysis and research systematically errors resulting from different schemes of spatialization of population data. In view of facts mentioned above. The author used national land cover map at scale of 1 to 250,000, national administrative division map and statistical population data at county level in 2005 as basic data to carry out the following researches:(1) Design of the spatialization schemes.(a)Design of scale of statistical samples: According to the relationship between population distribution and land cover types, selected the area of land cover at level I including forest, grassland, farmland, urban settlements, rural settlements, wetland & water, desert as dependent variables of a model, used the population data as independent variables of the model, established two models of population to types of land cover through multi-variables linear regression with statistical samples at county and prefecture levels respectively.(b) Design of addition or reduction of independent variables: With the area of farmland, urban settlements, and rural settlements, which have close relationship with population distribution, as mandatory independent variables, the other parameters including the area of forest, grassland, wetland & water and desert as optional independent variables, establish 28 models with statistical samples at county and prefecture levels respectively.(2) Index and method of errors evaluation. Selected correlation coefficients, fitted values, relative errors and absolute errors as evaluation index of errors for spatialization models and spatialization outputs. Besides indicators in the above, the ratio of amount of grid cells with negative value in spatialization outputs to entire study area was also considered as an evaluation index of errors for spatialization outputs.(3) Analysis of spatialization errors.(a) Compared two spatialization schemes with county-level and prefectural-level statistical population data as analysis samples respectively. It was concluded that the scheme of spatialization based on county-level samples is better with a correlation coefficient of 0.797 and a relative error of 6.2%.(b) Compared 28 spatialization schemes based on different parameters at county and prefecture levels. A conclusion was drawn that the spatialization schemes with farmland, urban settlements, rural settlements, wetland & water as dependent variables and county-level statistical data as analysis samples was best, the correlation coefficient of which was 0.797 and relative error 8.3%.(4) Selection and optimization of spatialization schemes. According to the result of analyzing error to select the best scheme, then it optimizes the scheme.(a) Optimized the spatialization scheme with county-level statistical data as analysis samples by zoning the whole county into 7 sub-regions. Optimized models of each sub-regions based on correlation coefficients, regression coefficients and scatter diagram that delete the abnormal value of scatter diagram. The optimized models had a correlation coefficient of greater than 0.92 for each sub-regions and a relative error of 0.134%.(b) Optimized the spatialization model with farmland, urban settlements, rural settlements, wetland & water as dependent variables and county-level statistical data as analysis samples. The numbers of parameter are kept unchanged when it was optimized. The model was optimized based on correlation coefficient, regression coefficient and scatter diagram that delete the abnormal value of scatter diagram. The optimized model had a correlation coefficient of 0.955 and a relative error of 0.131%.(c) A conclusion was drawn from analyzing and comparing spatialization results of two spatialization models(a) and(b) in the above that the spatialization model with farmland, urban settlements, rural settlements and wetland &water body as dependent variables and county-level statistical data as analysis samples was optimal scheme.In this paper, the author established an optimal scheme of population statistical data spatialization through analysis and comparison of different schemes. The scheme can express actual distribution of population in China. Therefore, the ideas, thoughts and final results from this paper will have a certain guidance and reference value on spatialization of statistical population data in the future.

【关键词】 人口统计数据空间化方案误差评价
【Key words】 PopulationStatistical dataSpatializationSchemeError evaluation
  • 【网络出版投稿人】 河南大学
  • 【网络出版年期】2016年 06期
  • 【分类号】P208;C924.2
  • 【被引频次】6
  • 【下载频次】437
节点文献中: 

本文链接的文献网络图示:

本文的引文网络