Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it. We use Transluce, an AI-driven mechanistic-interpretability platform, to examine LLaMA 3.1-8B-Instruct on 54 moral prompts in four batteries: 17 dilemmas, policy, and meta-ethical questions (B1); 6 role-playing scenarios (B3); and a controlled trolley contrast varying the switching mechanism with people fixed (B4, 15 prompts) or identity attributes with mechanism fixed (B5, 16 prompts). Two complementary metric families, five cluster-level metrics and a six-metric neuron-level panel, converge on a Situational Anchor Effect: domain-specific representations dominate the top of the activation list across every battery. The model's ethics-labeled capacity stays essentially constant; its salience (rank, priority, top-of-list presence) is highly sensitive to the interpretive frame the prompt selects. The B4-vs-B5 contrast confirms the model attends to whichever surface feature varies: aggregate ethics metrics are indistinguishable, but the dominant non-ethics distractor mirrors the design. A multi-temperature audit identifies a candidate ethics neuron (L16/N3837) stable across temperatures; a cross-model behavioral proxy on two frontier models yields preliminary evidence of divergence in self-reported moral focus, consistent with an Alignment Wrapper in which RLHF re-orders surface text without removing underlying domain-first frames. We unify these as Frame-Conditioned Moral Computation: the prompt's surface vocabulary selects a feature manifold, and the moral conclusion is downstream of that selection. Behavioral alignment must be supplemented by Mechanistic Alignment: a research program asking whether ethics-related features can be shown causally privileged under controlled frame variation, not merely loud in the explanation.


翻译:大语言模型在道德提示上的行为审计仅衡量模型的输出文本,而非其内部产生该输出的计算过程。我们使用人工智能驱动的机制可解释性平台Transluce,对LLaMA 3.1-8B-Instruct进行了54个道德提示的审计,涵盖四个测试组:17个困境、政策与元伦理问题(B1组);6个角色扮演场景(B3组);以及两个对照测试组——开关机制不同但人员固定的有轨电车变体问题(B4组,15个提示)和人员身份属性不同但机制固定的变体问题(B5组,16个提示)。两类互补的度量族——五个聚类级指标与六个神经元的度量面板——共同揭示了一种"情景锚定效应":每个测试组中激活列表顶部的表征均由领域特定表示主导。模型的伦理标注能力基本保持恒定;但其显著性(排名、优先级、列表顶部出现频率)高度依赖于提示所选定的解释性框架。B4组与B5组的对比证实,模型关注的是两者间变化的表面特征:聚合层面的伦理度量无法区分,但主导的非伦理干扰项镜像反映了实验设计的变化。一项多温度审计识别出一个跨温度稳定的候选伦理神经元(第16层/第3837号神经元);对两个前沿模型的跨模型行为代理测试提供了自我报告道德焦点差异的初步证据,这与"对齐包装"假说一致——即RLHF重排了表层文本,但未消除底层的领域优先框架。我们将这些发现统一为"帧条件化的道德计算":提示的表层词汇选择了一个特征流形,而道德结论是该选择的下游产物。行为对齐必须辅以"机制对齐":这是一项研究议程,旨在探明在受控框架变化下,伦理相关特征是否能够在因果关系上被证明具有优先性,而不仅仅是在解释中表现突出。

0
下载
关闭预览

相关内容

Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
迈向可信的人工智能:伦理和稳健的大型语言模型综述
专知会员服务
40+阅读 · 2024年7月28日
如何检测LLM内容?UCSB等最新首篇《LLM生成内容检测》综述
大模型道德价值观对齐问题剖析
专知会员服务
79+阅读 · 2023年10月3日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
NLP 与 NLU:从语言理解到语言处理
AI研习社
15+阅读 · 2019年5月29日
深入理解BERT Transformer ,不仅仅是注意力机制
大数据文摘
22+阅读 · 2019年3月19日
【泡泡图灵智库】密集相关的自监督视觉描述学习(RAL)
泡泡机器人SLAM
11+阅读 · 2018年10月6日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
5+阅读 · 9月3日
何为协作武器?
专知会员服务
9+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
13+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
8+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
《美陆军野战手册(2026年):特种部队》
专知会员服务
8+阅读 · 8月31日
相关VIP内容
Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
迈向可信的人工智能:伦理和稳健的大型语言模型综述
专知会员服务
40+阅读 · 2024年7月28日
如何检测LLM内容?UCSB等最新首篇《LLM生成内容检测》综述
大模型道德价值观对齐问题剖析
专知会员服务
79+阅读 · 2023年10月3日
相关基金
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员