Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. This abstraction gives practitioners little visibility into what their data actually teaches models, allowing spurious correlations to be learned by a model and inducing undesirable behaviors such as over-stylization and sycophancy. To address this problem, we ask: can we inspect a preference dataset before optimization and decide, at the level of concepts, which behaviors a model should be allowed to learn? Motivated by this, we introduce a data-centric post-training pipeline that uses interpretability protocols to develop statistical hypotheses for the latent concepts separating preferred from dispreferred generations, making them explicit for fine-grained user feedback. Building on this view, we unify several interpretability-based training protocols as ways of shaping rewards via feature or data interventions. Empirically, we show that our pipeline diagnoses undesirable signals in existing preference data, mitigates off-target learning, and can also help amplify or shape desired properties such as safeguards and model personality. More broadly, our results suggest that interpretability can turn post-training from optimizing opaque proxy rewards into a process of auditing and sculpting the learning signal itself.
翻译:语言模型的后训练是塑造模型行为的主要阶段,但其本质上仍涉及对汇总多元目标的标量奖励进行优化。这种抽象化处理使实践者难以洞察数据实际教授模型的内容,导致模型习得虚假相关性,并引发过度风格化、谄媚性等不良行为。针对这一问题,我们提出疑问:能否在优化前审视偏好数据集,从概念层面判定模型应被允许学习哪些行为?受此启发,我们提出一种以数据为中心的后训练流程,该流程利用可解释性协议对区分偏好与非偏好生成内容的潜在概念形成统计假设,并将其显式化以供细粒度用户反馈。基于这一视角,我们将多种基于可解释性的训练协议统一为通过特征或数据干预塑造奖励的方式。实验表明,我们的流程能够诊断现有偏好数据中的不良信号,缓解非目标学习,并有助于增强或塑造特定属性(如安全防护与模型人格)。更广泛而言,我们的研究结果表明,可解释性可将后训练从优化不透明代理奖励的过程,转变为审计与塑造学习信号本身的过程。