节点文献

从人类偏好到自主对齐:大语言模型对齐方法综述

From Human Preference to Autonomous Alignment: A Survey of Large Language Model Alignment

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 窦士涵张明黄萱菁柳世纯沈钰炯张家政黄宸颢陈佳逸郑惠元周玮康桂韬张奇

【Author】 DOU Shihan;ZHANG Ming;HUANG Xuanjing;LIU Shichun;SHEN Yujiong;ZHANG Jiazheng;HUANG Chenhao;CHEN Jiayi;ZHENG Huiyuan;ZHOU Weikang;GUI Tao;ZHANG Qi;Computation and Artificial Intelligence Innovative College, Fudan University;Institute of Trustworthy Embodied Artificial Intelligence, Fudan University;Institute of Modern Languages and Linguistics, Fudan University;

【通讯作者】 黄萱菁;

【机构】 复旦大学计算与智能创新学院复旦大学可信具身智能研究院复旦大学现代语言学研究院

【摘要】 大语言模型对齐技术旨在确保模型在能力、行为和价值观方面与人类的长远利益保持一致。该文系统且全面地回顾了大语言模型对齐技术的发展历程,从全新的视角对这些技术进行了整理和分类,并将其发展脉络总结为三大类别:从人类偏好中模仿学习,从反馈信号中归纳学习,以及通过思考和沟通实现自主对齐。针对每一项技术的特点、优势和挑战,该文进行了详细阐述和总结。同时,该文还概述了用于评估大模型对齐技术表现的评测方法,讨论了当前大语言模型对齐技术所面临的挑战,并探讨了未来实现更完善对齐技术的可能发展方向,以推动对齐技术的进一步发展。

【Abstract】 Large Language Model(LLM) alignment aims to ensure that models align with human long-term interests in terms of capabilities, behaviors, and values. This survey offers a systematic and comprehensive review of the development of LLM alignment techniques, providing a novel framework to organize and classify these approaches. We categorize the development trajectory of LLM alignment into three primary domains: imitation learning from human preference, inductive learning from feedback signal, and autonomous alignment through reflection and communication. Each category is examined regarding its characteristics, advantages, and challenges. Moreover, we also summarize the methods used to evaluate the performance of LLM alignment techniques and discuss potential future directions to further advance the field.

【基金】 广东省重点研发计划(2024B0101050003);国家自然科学基金(62441602)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2025年10期
  • 【分类号】TP18
  • 【下载频次】113
节点文献中: 

本文链接的文献网络图示:

本文的引文网络