We present a novel study analyzing the effects of various prompt loss token weights (PLW) for supervised instruction fine-tuning (SIFT). While prompt-masking (PLW = 0) is common for SIFT, some fine-tuning APIs support fractional PLWs and suggest that using a small non-zero PLW can help stabilize learning when fine-tuning on short-completion data. However, there has never been a study confirming this claim, and OpenAI, a major cloud-based SIFT provider, recently removed this parameter from their fine-tuning API. We found that performance of models fine-tuned on short-completion data had a statistically-significant negative quadratic relationship with PLW. Using small values (0.01 - 0.5) of PLW produced better results on multiple-choice and short-generation benchmarks (outperforming models fine-tuned on long-completion data) while large values (~ 1.0) of PLW produced better results on long-generation benchmarks. We explained this effect and verified its importance through additional experiments. This research serves as a warning to API providers about the importance of providing a PLW parameter for SIFT.
翻译:本研究首次系统分析了监督式指令微调中不同提示损失令牌权重的影响机制。尽管提示掩码(PLW = 0)在监督式指令微调中普遍应用,部分微调接口支持分数权重参数,并建议在短补全数据微调时采用较小非零权重以稳定学习过程。然而该主张始终缺乏实证研究支持,而作为主流云端监督式指令微调服务商,OpenAI近期已从其微调接口移除此参数。实验发现:在短补全数据微调时,模型性能与提示损失权重呈现统计学显著的负二次相关关系。采用较小权重值(0.01-0.5)能在多项选择与短文本生成基准测试中获得更优表现(超越长补全数据微调模型),而较大权重值(约1.0)则在长文本生成任务中表现更佳。我们通过理论阐释与补充实验验证了该效应的重要性。本研究为API服务商提供了重要警示:保留提示损失权重参数对监督式指令微调具有关键意义。