AI alignment work is important from both a commercial and a safety lens. With this paper, we aim to help actors who support alignment efforts to make these efforts as effective as possible, and to avoid potential adverse effects. We begin by suggesting that institutions that are trying to act in the public interest (such as governments) should aim to support specifically alignment work that reduces accident or misuse risks. We then describe four problems which might cause alignment efforts to be counterproductive, increasing large-scale AI risks. We suggest mitigations for each problem. Finally, we make a broader recommendation that institutions trying to act in the public interest should think systematically about how to make their alignment efforts as effective, and as likely to be beneficial, as possible.
翻译:AI对齐工作从商业和安全角度来看都至关重要。本文旨在帮助支持对齐工作的参与者尽可能有效地开展这些工作,并避免潜在的不利影响。我们首先建议,努力维护公共利益的机构(如政府)应重点支持那些能够降低事故或滥用风险的特定对齐工作。随后,我们描述了可能导致对齐工作适得其反、增加大规模AI风险的四个问题,并针对每个问题提出了缓解措施。最后,我们提出更广泛的建议:旨在维护公共利益的机构应系统性思考如何使其对齐工作尽可能高效且有益。