节点文献

基于虚实迁移强化学习的机器人按钮操作策略研究

Research on Robot Button Operation Policy Based on Sim-to-Real Reinforcement Learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 龙晖午肖聚亮赵炜刘海涛朱林陈斌

【Author】 Long Huiwu;Xiao Juliang;Zhao Wei;Liu Haitao;Zhu Lin;Chen Bin;Key laboratory of Mechanism Theory and Equipment Design of Ministry of Education,Tianjin University;Beijing Spacecraft Manufacturing Co.,Ltd.,China Academy of Space Technology;Corporate Technology China,Siemens Ltd.;

【通讯作者】 肖聚亮;

【机构】 天津大学机构理论与装备设计教育部重点实验室中国空间技术研究院北京卫星制造厂有限公司西门子中国研究院

【摘要】 具身智能概念的快速发展对智能体与物理世界的交互能力提出了更高要求.在以机器人为代表的智能载体与环境的交互过程中,主要依靠力反馈信号以决定其动作输出的任务称为力交互任务,例如零件装配、按钮操作和门窗开合等.针对此类任务交互对象种类繁多、力反馈特性各不相同的挑战,提出了一种具身智能训练方法,基于虚实迁移(sim-to-real)的概念和强化学习方法搭建了机器人高级力交互操作技能学习训练框架,赋予了机器人安全、准确、适应性强大的力交互操作能力.以经典的机器人力交互场景——按钮操作任务为例:首先,基于域随机化的方法在虚拟环境中构建了大量按钮模型,并从接触刚度的角度划分了机器人与按钮之间的接触阶段;然后,模仿人类在按钮操作过程中的感知方式,结合在线刚度估计算法,在虚拟环境下使用近端策略优化(PPO)算法训练机器人的按钮操作技能;最后,通过sim-to-real方法将得到的预训练策略直接部署在真实机器人上,在策略迁移后的实机操作实验中得到了良好的结果.在针对具有不同力反馈特性的按钮进行操作的泛化能力测试实验中,经上述方法训练得到的策略展现了远优于现有方法的泛化性能.

【Abstract】 The rapid development of the concept of embodied intelligence has increased the requirements for the interaction capabilities of agents with the physical world. During the interaction between intelligent carriers,such as robots and the environment,tasks that rely primarily on force feedback signals to determine action outputs are known as force interaction tasks,including component assembly,button operation,and door/window manipulation. To address the challenges posed by the diversity of interaction objects and the varying force feedback characteristics during such tasks,a training method for embodied intelligence was proposed. A training framework for learning advanced robotic force interaction skills was developed based on the sim-to-real concept,as well as reinforcement learning,thereby enabling robots to safely,accurately,and adaptively perform force interaction tasks. Considering the classic robotic force interaction scenario??button operation??as an example,numerous button models were constructed in a virtual environment using domain randomization,and the contact phases between the robot and the button were categorized based on contact stiffness. Thereafter,inspired by human perception of button operation,an online stiffness estimation algorithm was incorporated,and the proximal policy optimization(PPO)algorithm was employed to train the button operation skills of the robot in the virtual environment. Finally,the pretrained policy was directly deployed onto a real robot through the sim-to-real method,achieving favorable results in real-world experiments. In generalization tests on buttons with different force feedback characteristics,the policy,which was trained using the proposed method,demonstrated significantly superior generalization performance compared with the existing approaches.

【基金】 国家自然科学基金资助项目(52175025,52325501)~~
  • 【文献出处】 天津大学学报(自然科学与工程技术版) ,Journal of Tianjin University(Science and Technology) , 编辑部邮箱 ,2026年04期
  • 【分类号】TP242;TP18
  • 【下载频次】19
节点文献中: 

本文链接的文献网络图示:

本文的引文网络