节点文献
电梯群控动态配置的强化学习简化算法
Reinforcement Learning Simplification Algorithm for Dynamic Dispatching of Elevator Group Control
【摘要】 根据马尔科夫决策过程和Q-learning算法,通过简化处理求解电梯群控系统在上行峰值期间的最优配置问题。给出电梯群控系统最优配梯的步骤、框图和量化计算求解例;提出对应的报酬R即目标函数所包括的5项分函数;给出在候梯期间和电梯运行期间的到达人数和到达概率公式及工程意义;指出候梯时间可以用"部分"运行周期即返回时间来表示。由此扩大了Basset公式的应用,为深入研究电梯群控系统打下了基础。
【Abstract】 According to Markov Decision Process and Q-learning algorithm, the optimal dispatching of elevator group control system during the performing up-peak is solved by simplified processing. The steps, block diagrams and quantitative calculation examples of optimal dispatching in elevator group control system are given. The corresponding reward R, which is the objective function, including the five sub-functions is proposed. During the waiting and the elevator running, the formulas of arrival number and probability of arrival and their engineering significance are given. It is pointed out that the waiting time can be expressed by the " partial" round-trip-time, which is the return time. Therefore, the application of Basset formula is expanded and a foundation for the in-depth study of elevator group control system is laid.
【Key words】 Markov Decision Process; reinforcement learning; reward function; number of possible stops;
- 【文献出处】 中国电梯 ,China Elevator , 编辑部邮箱 ,2020年02期
- 【分类号】TU857
- 【被引频次】2
- 【下载频次】132