节点文献
从人类偏好到自主对齐:大语言模型对齐方法综述
From Human Preference to Autonomous Alignment: A Survey of Large Language Model Alignment
【摘要】 大语言模型对齐技术旨在确保模型在能力、行为和价值观方面与人类的长远利益保持一致。该文系统且全面地回顾了大语言模型对齐技术的发展历程,从全新的视角对这些技术进行了整理和分类,并将其发展脉络总结为三大类别:从人类偏好中模仿学习,从反馈信号中归纳学习,以及通过思考和沟通实现自主对齐。针对每一项技术的特点、优势和挑战,该文进行了详细阐述和总结。同时,该文还概述了用于评估大模型对齐技术表现的评测方法,讨论了当前大语言模型对齐技术所面临的挑战,并探讨了未来实现更完善对齐技术的可能发展方向,以推动对齐技术的进一步发展。
【Abstract】 Large Language Model(LLM) alignment aims to ensure that models align with human long-term interests in terms of capabilities, behaviors, and values. This survey offers a systematic and comprehensive review of the development of LLM alignment techniques, providing a novel framework to organize and classify these approaches. We categorize the development trajectory of LLM alignment into three primary domains: imitation learning from human preference, inductive learning from feedback signal, and autonomous alignment through reflection and communication. Each category is examined regarding its characteristics, advantages, and challenges. Moreover, we also summarize the methods used to evaluate the performance of LLM alignment techniques and discuss potential future directions to further advance the field.
【Key words】 artificial intelligence; large language model; value and capability alignment; reinforcement learning;
- 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2025年10期
- 【分类号】TP18
- 【下载频次】113