节点文献
非平稳MDP平均模型及其滚动式算法
AVERAGE MODEL IN NONHOMOGENEOUS MARKOV DECISION PROCESSES AND ROLLING HORIZON ALGORITHM
【摘要】 本文考虑可数状态空间非平稳马尔可夫决策过程(MDP)的平均目标.首先,我们指出并改正了Park,et,al[1]和Alden,etal[2]的错误,并在弱于Park,etal[1]的条件下,借助于新建立的最优方程,证明了最优平均值的收敛性和平均最优马氏策略的存在性.其次,给出了ε(>0)-平均最优马氏策略的滚动式算法.
【Abstract】 in this paper, we consider denumerable state nonstationary Markov decisionprocesses with average criterion. First, we point out and correct some mistakes in Park, et al.~[1]and in Alden, et al.~[2]. By the optimal equation built in this paper, we prove the existence ofε(≥ 0)-optimal Markov policies for average criterion and the convergence of optimal averagevalue under conditions which are weaker than those used Park, et al.~[1]. Secondly, we providea rolling horizon algorithm for ε(> 0)-optimal policies.
【关键词】 非平稳MDP;
平均目标;
ε(≥0)-平均最优马氏策略;
滚动式算法;
最优方程;
【Key words】 Nonhomogeneous Markov decision processes; average criterion; ε(≥0 )-optimal policies; optimal equation; rolling horizon algorithm.;
【Key words】 Nonhomogeneous Markov decision processes; average criterion; ε(≥0 )-optimal policies; optimal equation; rolling horizon algorithm.;
【基金】 国家青年基金;国家自然科学基金;广东省自然科学基金;亚太运筹中心资助
- 【文献出处】 系统科学与数学 ,JOURNAL OF SYSTEMS SCIENCE AND MATHEMATICAL SCIENCES , 编辑部邮箱 ,1999年04期
- 【分类号】O224
- 【被引频次】4
- 【下载频次】75