As large language models (LLMs) grow more powerful, concerns around potential harms like toxicity, unfairness, and hallucination threaten user trust. Ensuring beneficial alignment of LLMs with human values through model alignment is thus critical yet challenging, requiring a deeper understanding of LLM behaviors and mechanisms. We propose opening the black box of LLMs through a framework of holistic interpretability encompassing complementary bottom-up and top-down perspectives. The bottom-up view, enabled by mechanistic interpretability, focuses on component functionalities and training dynamics. The top-down view utilizes representation engineering to analyze behaviors through hidden representations. In this paper, we review the landscape around mechanistic interpretability and representation engineering, summarizing approaches, discussing limitations and applications, and outlining future challenges in using these techniques to achieve ethical, honest, and reliable reasoning aligned with human values.
翻译:随着大型语言模型(LLMs)能力的不断增强,毒性、不公平性和幻觉等潜在危害引发的担忧威胁着用户信任。通过模型对齐确保LLMs与人类价值观的有益对齐至关重要且充满挑战,这需要更深入地理解LLM的行为与机制。我们提出通过一个涵盖自下而上与自上而下互补视角的整体可解释性框架来打开LLM的黑箱。自下而上的视角借助机械可解释性,聚焦于组件功能与训练动态;自上而下的视角则利用表征工程,通过隐藏表征来分析行为。本文围绕机械可解释性与表征工程的相关研究进行综述,总结现有方法,探讨其局限性与应用,并勾勒出利用这些技术实现符合人类价值观的伦理、诚实与可靠推理所面临的未来挑战。