Whether language models can systematically generalize remains actively debated. Yet empirical performance is jointly shaped by multiple factors such as training data, training paradigms, and inference-time strategies, making failures difficult to interpret. We introduce a controlled synthetic environment based on shortest-path planning, a canonical composable sequential optimization problem. The setup enables clean separation of these factors and supports two orthogonal axes of generalization: spatial transfer to unseen maps and length scaling to longer-horizon problems. We find that models exhibit strong spatial transfer but consistently fail under length scaling due to recursive instability. We further analyze how distinct stages of the learning pipeline influence systematic problem-solving: for example, data coverage sets capability limits; reinforcement learning improves training stability but does not expand those limits; and inference-time scaling enhances performance but cannot rescue length-scaling failures.
翻译:语言模型能否实现系统性泛化仍是一个备受争议的问题。然而,实证性能受到训练数据、训练范式、推理时策略等多重因素的共同影响,这使得失败原因难以解释。我们基于最短路径规划——一个经典的组合顺序优化问题——构建了一个受控的合成实验环境。该实验设置能够清晰分离上述各因素,并支持两个正交的泛化维度:面向未见地图的空间迁移,以及面向更长路径问题的长度缩放。我们发现模型展现出强大的空间迁移能力,但由于递归不稳定性,在长度缩放任务中持续表现不佳。我们进一步分析了学习流程的不同阶段如何影响系统性问题的求解能力:例如,数据覆盖范围决定了能力上限;强化学习提升了训练稳定性但未能突破该上限;而推理时计算量扩展虽能提升性能,却无法挽救长度缩放中的失败。