We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reasoning, we develop and publicly release the Animal Harm Benchmark (AHB), a 26-question evaluation spanning 13 ethical dimensions, publicly available as a dataset and Inspect evaluation. On the AHB, training with 3000 documents achieves 77% compared to 40% for instruction-tuning approaches, with generalization to human compassion and no degradation in standard safety benchmarks or capabilities. However, subsequent unrelated instruction-tuning degrades the intervention, with the advantage disappearing after 5000 samples. Our exploratory results suggest document-based value interventions may require explicit preservation strategies to remain effective through typical training pipelines.
翻译:我们研究了基于合成文档的微调对价值对齐的稳健性,以动物同情心作为一项本身具有重要意义且与现有对齐工作正交的价值进行评估。为评估具有同情心的推理能力,我们开发并公开了动物伤害基准测试(AHB),该基准包含26道问题,涵盖13个伦理维度,以数据集和Inspect评估形式公开。在AHB上,使用3000份文档进行训练可获得77%的准确率,而指令微调方法仅为40%,且该结果可泛化至人类同情心,同时标准安全基准或能力未出现退化。然而,后续不相关的指令微调会削弱该干预效果,5000个样本后优势完全消失。我们的探索性结果表明,基于文档的价值干预可能需要明确的保留策略,才能在典型训练流程中持续有效。