Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an external judge's multi-attribute quality scores on open-ended responses well above chance across three benchmarks. We introduce Self-Evaluation Elicitation (SEE), a method that surfaces this latent ability through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while leaving the answer untouched. From 160 unique examples, roughly 31x fewer than a reinforcement learning baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality. The elicited self-evaluation is sharply localized within the model's own token distribution and stable across judges it was never trained against, indicating a transferable notion of quality rather than a single judge's preference. These results reframe judge-aligned self-evaluation as a problem of elicitation rather than acquisition.
翻译:大型语言模型日益被其他模型评估,这引发了一个自然问题:模型能否预测评判者如何对其自身输出进行评分?我们发现,在针对性训练之前,该能力已普遍存在:通过少量示例提示,基础模型在三个基准测试中,对开放性回答的多属性质量评分预测已显著高于随机水平。我们引入自我评估激发(SEE)方法,该方法通过短周期来激活这一潜在能力,周期包含一个校准耦合的强化学习阶段,用于改进答案并预测评判者,随后是一个掩码蒸馏阶段,在不改动答案的前提下增强预测精度。仅需160个独特示例(约为强化学习基线所需数据的31分之一),SEE在三个基准测试中提升了留出测试的校准性能,同时保持了答案质量。所激发的自我评估严格局限于模型自身的词元分布,且对未经训练的其他评判者保持稳定,这表明其反映的是可迁移的质量概念,而非单一评判者的偏好。这些结果将评判者对齐的自我评估重新定义为激发问题而非习得问题。