Training large language models (LLMs) with open-domain instruction data has yielded remarkable success in aligning to end tasks and user preferences. Extensive research has highlighted that enhancing the quality and diversity of instruction data consistently improves performance. However, the impact of data complexity, as a crucial metric, remains relatively unexplored in three aspects: (1) scaling law, where the sustainability of performance improvements with increasing complexity is uncertain, (2) additional tokens, whether the improvement brought by complexity comes from introducing more training tokens, and (3) curriculum tuning, where the potential advantages of incorporating instructions ranging from easy to difficult are not yet fully understood. In this paper, we propose \textit{tree-instruct} to systematically enhance the complexity of instruction data in a controllable manner. This approach adds a specified number of nodes into the instruction semantic tree, yielding new instruction data based on the modified tree. By adjusting the number of added nodes, we can control the difficulty level in the modified instruction data. Our preliminary experiments reveal the following insights: (1) Increasing complexity consistently leads to sustained performance improvements. For instance, using 1,000 instruction data and 10 nodes resulted in a substantial 24\% increase in win rate. (2) Under the same token budget, a few complex instructions outperform diverse yet simple instructions. (3) Curriculum instruction tuning might not yield the anticipated results; focusing on increasing complexity appears to be the key.
翻译:训练大型语言模型(LLMs)使用开放域指令数据已在与最终任务及用户偏好对齐方面取得了显著成功。大量研究强调,提升指令数据的质量和多样性能够持续改善模型性能。然而,数据复杂性作为关键指标,其影响在以下三个方面尚未充分探索:(1)缩放定律——性能提升随复杂性增加的可持续性尚不确定;(2)额外token——复杂性带来的改进是否源于引入更多训练token;(3)课程调优——从简单到困难指令的渐进式训练可能带来的优势尚未完全明晰。本文提出\textit{tree-instruct}方法,以可控方式系统性增强指令数据的复杂性。该方法通过向指令语义树中添加指定数量的节点,基于修改后的树生成新指令数据。通过调整添加节点数量,可控制修改后指令数据的难度级别。初步实验揭示了以下洞察:(1)增加复杂性可持续带来性能提升。例如,使用1000条指令数据和10个节点,胜率提升了24%。(2)在相同token预算下,少量复杂指令优于多样但简单的指令。(3)课程式指令调优可能未产生预期效果;聚焦于增加复杂性似乎才是关键。