The ML literature contains many distinct concepts falling under the heading of 'AI alignment'. After noting three concepts of AI alignment in the context of their corresponding research programs, we claim that realistic interventions may promote 'AI alignment' under one conception while being actively counterproductive from the perspective of others. We suggest that tensions between alignment ideals emerge due to differences in background threat-models, alongside differences in normative orientations. In light of our analysis, researchers aiming to further the goal of 'AI alignment' should do five things. First, they should not conflate distinctions of policy and distinctions of scientific scope; second, methodological disagreements should be acknowledged explicitly; third, researchers should distinguish between 'AI alignment' as a high-level ideal and specific 'alignment proxies' used in empirical research; fourth, they should use more granular concepts to identify both the source and nature of possible AI harms/benefits; fifth, they should explicitly acknowledge the diversity of 'alignment' concepts in both empirical work and in communication with non-technical audiences.
翻译:机器学习文献中存在多种不同的概念,均被归入“AI对齐”这一标题下。在注意到与其相应研究计划相关的三个“AI对齐”概念后,我们声称:在某种概念下,现实干预措施可能促进“AI对齐”,但从其他概念视角看,这些措施可能适得其反。我们认为,对齐理想之间的张力源于背景威胁模型差异以及规范取向差异。基于我们的分析,旨在推进“AI对齐”目标的研究者应做到五点:第一,不应混淆政策区分与科学范围区分;第二,应明确承认方法论分歧;第三,应将“AI对齐”作为高层理想与实证研究中使用的具体“对齐代理指标”加以区分;第四,应使用更细粒度的概念来识别可能的人工智能危害/收益的来源与性质;第五,在实证工作及与非技术受众的交流中,应明确承认“对齐”概念的多样性。