An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessive conservatism discourages exploration. How much behavior change is too much? We show how to use any safe reference policy as a probabilistic regulator for any optimized but untested policy. Conformal calibration on data from the safe policy determines how aggressively the new policy can act, while provably enforcing the user's declared risk tolerance. Unlike conservative optimization methods, we do not assume the user has identified the correct model class nor tuned any hyperparameters. Unlike previous conformal methods, our theory provides finite-sample guarantees even for non-monotonic bounded loss functions. Our experiments on applications ranging from natural language question answering to biomolecular engineering show that safe exploration is not only possible from the first moment of deployment, but can also improve performance.
翻译:代理必须尝试新行为以进行探索和改进。在高风险环境中,违反安全约束的代理可能造成伤害,且必须被下线处理,从而终止任何未来交互。模仿旧行为是安全的,但过度保守会抑制探索。行为改变多少才算过度?我们展示了如何利用任何安全参考策略作为概率调节器,以约束任何已优化但未经测试的策略。基于安全策略数据的保形校准决定了新策略可以采取行动的激进程度,同时可证明地强制执行用户声明的风险容忍度。与保守优化方法不同,我们既不假设用户已确定正确的模型类别,也不假设其已调整任何超参数。与以往的保形方法不同,我们的理论即使对于非单调有界损失函数也能提供有限样本保证。我们在从自然语言问答到生物分子工程等应用上的实验表明,安全探索不仅可以在部署初始阶段实现,还能提升性能。