In this work, we probe the ability of a language model to demonstrate spatial reasoning from unstructured text, mimicking human capabilities and automating a process that benefits many downstream media applications. Concretely, we study the narrative-to-play task: inferring stage-play layouts (scenes, speaker positions, movements, and room types) from text that lacks explicit spatial, positional, or relational cues. We then introduce a dramaturgy-inspired deterministic evaluation suite and, finally, a training and inference recipe that combines rejection SFT using Best-of-N sampling with RL from verifiable rewards via GRPO. Experiments on a text-only corpus of classical English literature demonstrate improvements over vanilla models across multiple metrics (character attribution, spatial plausibility, and movement economy), as well as alignment with an LLM-as-a-judge and subjective human preferences.
翻译:在这项工作中,我们探究语言模型从非结构化文本中展现空间推理能力,从而模拟人类能力并自动化一个有益于众多下游媒体应用的过程。具体而言,我们研究叙事到剧本的任务:从缺乏显式空间、位置或关系线索的文本中推断舞台布局(场景、说话者位置、移动轨迹及房间类型)。随后我们引入一套受戏剧学启发的确定性评估基准,并最终提出一种结合最佳N采样(Best-of-N)的拒绝式微调(SFT)与基于可验证奖励的GRPO强化学习(RL)的训练与推理方案。在古典英语文学纯文本语料上的实验表明,与基础模型相比,该方案在多个指标(人物归属、空间合理性及移动经济性)上均有改进,同时与采用LLM作为评判标准的方法及人类主观偏好保持一致性。