Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it. We use Transluce, an AI-driven mechanistic-interpretability platform, to examine LLaMA 3.1-8B-Instruct on 54 moral prompts in four batteries: 17 dilemmas, policy, and meta-ethical questions (B1); 6 role-playing scenarios (B3); and a controlled trolley contrast varying the switching mechanism with people fixed (B4, 15 prompts) or identity attributes with mechanism fixed (B5, 16 prompts). Two complementary metric families, five cluster-level metrics and a six-metric neuron-level panel, converge on a Situational Anchor Effect: domain-specific representations dominate the top of the activation list across every battery. The model's ethics-labeled capacity stays essentially constant; its salience (rank, priority, top-of-list presence) is highly sensitive to the interpretive frame the prompt selects. The B4-vs-B5 contrast confirms the model attends to whichever surface feature varies: aggregate ethics metrics are indistinguishable, but the dominant non-ethics distractor mirrors the design. A multi-temperature audit identifies a candidate ethics neuron (L16/N3837) stable across temperatures; a cross-model behavioral proxy on two frontier models yields preliminary evidence of divergence in self-reported moral focus, consistent with an Alignment Wrapper in which RLHF re-orders surface text without removing underlying domain-first frames. We unify these as Frame-Conditioned Moral Computation: the prompt's surface vocabulary selects a feature manifold, and the moral conclusion is downstream of that selection. Behavioral alignment must be supplemented by Mechanistic Alignment: a research program asking whether ethics-related features can be shown causally privileged under controlled frame variation, not merely loud in the explanation.
翻译:大语言模型在道德提示上的行为审计仅衡量模型的输出文本,而非其内部产生该输出的计算过程。我们使用人工智能驱动的机制可解释性平台Transluce,对LLaMA 3.1-8B-Instruct进行了54个道德提示的审计,涵盖四个测试组:17个困境、政策与元伦理问题(B1组);6个角色扮演场景(B3组);以及两个对照测试组——开关机制不同但人员固定的有轨电车变体问题(B4组,15个提示)和人员身份属性不同但机制固定的变体问题(B5组,16个提示)。两类互补的度量族——五个聚类级指标与六个神经元的度量面板——共同揭示了一种"情景锚定效应":每个测试组中激活列表顶部的表征均由领域特定表示主导。模型的伦理标注能力基本保持恒定;但其显著性(排名、优先级、列表顶部出现频率)高度依赖于提示所选定的解释性框架。B4组与B5组的对比证实,模型关注的是两者间变化的表面特征:聚合层面的伦理度量无法区分,但主导的非伦理干扰项镜像反映了实验设计的变化。一项多温度审计识别出一个跨温度稳定的候选伦理神经元(第16层/第3837号神经元);对两个前沿模型的跨模型行为代理测试提供了自我报告道德焦点差异的初步证据,这与"对齐包装"假说一致——即RLHF重排了表层文本,但未消除底层的领域优先框架。我们将这些发现统一为"帧条件化的道德计算":提示的表层词汇选择了一个特征流形,而道德结论是该选择的下游产物。行为对齐必须辅以"机制对齐":这是一项研究议程,旨在探明在受控框架变化下,伦理相关特征是否能够在因果关系上被证明具有优先性,而不仅仅是在解释中表现突出。