Augmented Markov decision process Q-Learning(AMDP-Q) was proposed,which was inspired by the thought of combination of AMDP, Monte Carlo-partially observable Markov decision process(MC-POMDP) and Q-learning.Firstly a lower-dimensional sufficient statistic was taken to represent the belief state space.In the common situation,a good choice was the tuple of the maximum likelihood state and the entropy of the belief.The new space composed of this tuple was referred to as the augmented state space.Secondly,a set ...