Our intention is to provide a definitive reference on what it would take to safely make use of generative/predictive models in the absence of a solution to the Eliciting Latent Knowledge problem. Furthermore, we believe that large language models can be understood as such predictive models of the world, and that such a conceptualization raises significant opportunities for their safe yet powerful use via carefully conditioning them to predict desirable outputs. Unfortunately, such approaches also raise a variety of potentially fatal safety problems, particularly surrounding situations where predictive models predict the output of other AI systems, potentially unbeknownst to us. There are numerous potential solutions to such problems, however, primarily via carefully conditioning models to predict the things we want (e.g. humans) rather than the things we don't (e.g. malign AIs). Furthermore, due to the simplicity of the prediction objective, we believe that predictive models present the easiest inner alignment problem that we are aware of. As a result, we think that conditioning approaches for predictive models represent the safest known way of eliciting human-level and slightly superhuman capabilities from large language models and other similar future models.
翻译:我们的目标是提供一个权威性参考,阐明在缺乏“引发潜在知识”问题解决方案的情况下,如何安全地使用生成/预测模型。此外,我们认为大型语言模型可被理解为这种对世界进行预测的模型,而这一概念化通过精心条件化模型以预测期望输出,为它们的安全且强大的应用带来了重大机遇。然而,这类方法也引发了多种潜在的致命安全问题,特别是当预测模型预测其他人工智能系统(可能在我们不知情的情况下)的输出时。对此类问题存在大量潜在解决方案,主要途径是精心条件化模型以预测我们期望的事物(例如人类),而非我们不期望的事物(例如恶意人工智能)。此外,由于预测目标的简单性,我们认为预测模型呈现了我们所知的最简单的内部对齐问题。因此,我们认为针对预测模型的条件化方法是当前从大型语言模型及其他类似未来模型中引发人类级及略超人类能力的最安全途径。