节点文献

CTMDP基于随机平稳策略的仿真优化算法(英文)

A Simulation Optimization Algorithm for CTMDPs Based on Randomized Stationary Policies

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 唐昊奚宏生殷保群

【Author】 TANG Hao XI Hong-Sheng YIN Bao-Qun(Department of Automation , University of Science and Technology of China , Hefei 230026) (Department of Computer, Hefei University of Technology , Hefei 230009)

【机构】 中国科学技术大学自动化系中国科学技术大学自动化系 合肥 230026 合肥工业大学计算机系 合肥 230009合肥 230026合肥 230026

【摘要】 基于Markov性能势理论和神经元动态规划(NDP)方法,研究一类连续时间Markov决策过程(MDP)在随机平稳策略下的仿真优化问题,给出的算法是把一个连续时间过程转换成其一致化Markov链,然后通过其单个样本轨道来估计平均代价性能指标关于策略参数的梯度,以寻找次优策略,该方法适合于解决大状态空间系统的性能优化问题。并给出了一个受控Markov过程的数值实例.

【Abstract】 Based on the theory of Markov performance potentials and neuro-dynamic programming (NDP) methodology, we study simulation optimization algorithm for a class of continuous time Markov decision processes (CTMDPs) under randomized stationary policies. The proposed algorithm will estimate the gradient of average cost performance measure with respect to policy parameters by transforming a continuous time Markov process into a uniform Markov chain and simulating a single sample path of the chain. The goal is to look for a suboptimal randomized stationary policy. The algorithm derived here can meet the needs of performance optimization of many difficult systems with large-scale state space. Finally, a numerical example for a controlled Markov process is provided.

【基金】 National Natural Science Foundation of P.R.China(60274012); the Natural Science Foundation of Anhui Province(01042308)
  • 【文献出处】 自动化学报 ,Acta Automatica Sinica , 编辑部邮箱 ,2004年02期
  • 【分类号】TP391.9
  • 【被引频次】3
  • 【下载频次】71
节点文献中: 

本文链接的文献网络图示:

本文的引文网络