Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Training (PST), a label-free fine-tuning framework in which multiple language models improve collaboratively by leveraging a cross-model aggregated response as an internal training signal. Given a prompt question, the models generate responses sequentially; the final aggregated answer, often more reliable than individual responses in practice, serves as an internal target for learning. We measure how informative each intermediate response is about the aggregate using pointwise mutual information (PMI), and use this signal to scale self-training updates. Responses already aligned with the aggregate are updated less, while less informative or misaligned responses are updated more. On mathematical reasoning benchmarks (SimulEq, Math500, and MultiArith), PST improves exact-match accuracy by 2.2 to 4.3 percentage points across Gemma-2-2B, LLaMA-3.2-1B, and Qwen-2.5-1.5B, and reduces the average generator-verifier gap (GV-Gap) by 26 to 40 percent, while requiring no external supervision or teacher-student hierarchy and relying solely on cross-model interactions. These results suggest that cross-model generations and peer-predictive feedback can serve as an effective approach for self-supervised training.
翻译:对于语言模型无需外部监督即可实现持续自我提升的机制仍是一个开放挑战。本文提出协同预测自训练(Peer-Predictive Self-Training, PST),一种无标签微调框架,其中多个语言模型通过利用跨模型聚合响应作为内部训练信号实现协作改进。给定提示问题后,各模型按序生成响应;最终聚合答案(实践中通常比个体响应更可靠)作为学习内部目标。我们采用点互信息(PMI)度量每个中间响应关于聚合答案的信息量,并基于此信号缩放自训练更新步长:与聚合答案一致的响应更新较少,而信息量不足或偏离的响应更新更多。在数学推理基准(SimulEq、Math500 和 MultiArith)上,PST 在 Gemma-2-2B、LLaMA-3.2-1B 和 Qwen-2.5-1.5B 模型中将精确匹配准确率提升 2.2 至 4.3 个百分点,同时将平均生成器-验证器差距(GV-Gap)降低 26% 至 40%,全程无需外部监督或师生层级结构,仅依赖跨模型交互。这些结果表明,跨模型生成与协同预测反馈可成为自监督训练的有效途径。