节点文献

非平稳MDP平均模型及其滚动式算法

AVERAGE MODEL IN NONHOMOGENEOUS MARKOV DECISION PROCESSES AND ROLLING HORIZON ALGORITHM

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 郭先平刘建庸刘克

【Author】 Guo Xianping(Department of Mathematics, Zhong Shan University; APORC 510275)Liu Jianyong ;Liu Ke(Institute Of Applied Mathematics, Academia Silica, Beijing 100080)

【机构】 中山大学数学系!亚太运筹中心广州510275中国科学院应用数学研究所!北京100080

【摘要】 本文考虑可数状态空间非平稳马尔可夫决策过程(MDP)的平均目标.首先,我们指出并改正了Park,et,al[1]和Alden,etal[2]的错误,并在弱于Park,etal[1]的条件下,借助于新建立的最优方程,证明了最优平均值的收敛性和平均最优马氏策略的存在性.其次,给出了ε(>0)-平均最优马氏策略的滚动式算法.

【Abstract】 in this paper, we consider denumerable state nonstationary Markov decisionprocesses with average criterion. First, we point out and correct some mistakes in Park, et al.~[1]and in Alden, et al.~[2]. By the optimal equation built in this paper, we prove the existence ofε(≥ 0)-optimal Markov policies for average criterion and the convergence of optimal averagevalue under conditions which are weaker than those used Park, et al.~[1]. Secondly, we providea rolling horizon algorithm for ε(> 0)-optimal policies.

【基金】 国家青年基金;国家自然科学基金;广东省自然科学基金;亚太运筹中心资助
  • 【文献出处】 系统科学与数学 ,JOURNAL OF SYSTEMS SCIENCE AND MATHEMATICAL SCIENCES , 编辑部邮箱 ,1999年04期
  • 【分类号】O224
  • 【被引频次】4
  • 【下载频次】75
节点文献中: 

本文链接的文献网络图示:

本文的引文网络