The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals during misaligned behaviour. The two primary toolkits are not jointly reported on the most alignment-relevant episodes. This note identifies two hypotheses that are qualitatively consistent with the published results: that the emotion vectors track functional emotions that causally drive behaviour, or that they are a projection of a richer situational-context structure onto human emotional axes. The hypotheses can be distinguished by cross-referencing the two toolkits on episodes where only one is currently reported: most directly, applying emotion probes to the strategic concealment episodes analysed only with SAE features. If emotion probes show flat activation while SAE features are strongly active, the alignment-relevant structure lies outside the emotion subspace. Which hypothesis is correct determines whether emotion-based monitoring will robustly detect dangerous model behaviour or systematically miss it.
翻译:Claude Mythos预览系统卡整合了情感向量、稀疏自编码器(SAE)特征及激活言语化器,用于研究模型在失调行为期间的内部状态。然而,这两类主要工具尚未在最相关的对齐行为情景中进行联合报告。本注释识别出两种与已发表结果定性一致的假设:其一,情感向量追踪因果驱动行为的"功能性情感";其二,情感向量是将更丰富的情境上下文结构投射到人类情感轴的结果。这两种假设可通过交叉引用目前仅报告单一工具集的情景加以区分——最直接的方法是将情感探针应用于仅使用SAE特征分析的策略性隐藏情景。若情感探针显示平坦激活而SAE特征高度活跃,则对齐相关结构位于情感子空间之外。哪种假设成立将决定基于情感的监控是能稳健检测危险模型行为,还是会系统性地遗漏此类行为。