We introduce StorySim, a programmable framework for synthetically generating stories to evaluate the theory of mind (ToM) and world modeling (WM) capabilities of large language models (LLMs). Unlike prior benchmarks that may suffer from contamination in pretraining data, or rely on an LLM for generation, StorySim produces novel, compositional story prompts anchored by a highly controllable Storyboard, enabling precise manipulation of character perspectives and events. Using StorySim, we evaluate LLMs across three kinds of ToM tasks: false belief, goal-directed, and preference attribution tasks. We then use our framework to design first- and second-order ToM tasks alongside WM tasks that control for the ability to track and model mental states. Our experiments across a suite of LLMs show that most models achieve higher accuracy on WM tasks than on ToM tasks, and that some models tend to reason more accurately when the subject of reasoning is a person rather than an inanimate object. Additionally, models struggle with goal-directed reasoning the most, and performance does not always scale with model size. All code for generating data and evaluations is freely available.
翻译:暂无翻译