节点文献

数据预处理对LSTM网络大气污染预测精度分析

Accuracy Analysis of Air Pollution Prediction for LSTM Network Based on Data Preprocessing

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 杜英魁张乙芳原忠虎关屏彭跃

【Author】 DU Yingkui;ZHANG Yifang;YUAN Zhonghu;GUAN Ping;PENG Yue;School of Information Engineering,Shenyang University;Shenyang Hengyuan Weiye Environmental Inspection Service Co.,Ltd.;Liaoning Environmental Monitoring Experimental Center;

【机构】 沈阳大学信息工程学院沈阳恒源伟业环境检测服务有限公司辽宁省环境监测实验中心

【摘要】 大气污染物浓度数据具有时序性和非线性的特点,针对时间序列数据中的异常值和缺失值问题,进行异常值和缺失值预处理对长短时记忆神经网络(LSTM)预测精度的影响分析。利用箱线图法判别数据序列中的异常值,以均值替换法、回归插补法和多重插补法进行缺失值的预处理,分别利用原始数据序列和不同预处理方法得到的数据序列,对多变量输入LSTM神经网络的大气污染物预测精度进行对比分析。实验结果表明,三种预处理方法均可明显改善LSTM模型的预测精度,多重插补法精度最高。

【Abstract】 Atmospheric pollutant concentration data is characterized by time series and non-linearity. For the problem of outliers and missing values in time series data,the influence of outlier and missing value preprocessing on the prediction accuracy of long and short time memory neural network(LSTM)is analyzed. The boxplot method is used to discriminate the outliers in the data sequence,and the mean value replacement method,the regression interpolation method and the multiple interpolation method are used to preprocess the missing values. By using the original data sequence and the data sequence obtained by different pretreatment methods,the prediction accuracy of air pollutants input into LSTM neural network is compared and analyzed. The experimental results show that the three data sequence preprocessing methods can significantly improve the prediction accuracy of the LSTM model and the multi-interpolation method has the highest accuracy.

【基金】 辽宁省重点研发计划指导计划项(编号:2018104013);辽宁省教育厅高等学校创新人才支持计划项目(编号:LR2016074);沈阳市中青年科技创新人才支持计划项目(编号:RC180338)资助
  • 【文献出处】 计算机与数字工程 ,Computer & Digital Engineering , 编辑部邮箱 ,2021年07期
  • 【分类号】X831;TP183
  • 【被引频次】2
  • 【下载频次】803
节点文献中: 

本文链接的文献网络图示:

本文的引文网络