We present a benchmark for assessing the capability of Large Language Models (LLMs) to discern intercardinal directions between geographic locations and apply it to three prominent LLMs: GPT-3.5, GPT-4, and Llama-2. This benchmark specifically evaluates whether LLMs exhibit a hierarchical spatial bias similar to humans, where judgments about individual locations' spatial relationships are influenced by the perceived relationships of the larger groups that contain them. To investigate this, we formulated 14 questions focusing on well-known American cities. Seven questions were designed to challenge the LLMs with scenarios potentially influenced by the orientation of larger geographical units, such as states or countries, while the remaining seven targeted locations less susceptible to such hierarchical categorization. Among the tested models, GPT-4 exhibited superior performance with 55.3% accuracy, followed by GPT-3.5 at 47.3%, and Llama-2 at 44.7%. The models showed significantly reduced accuracy on tasks with suspected hierarchical bias. For example, GPT-4's accuracy dropped to 32.9% on these tasks, compared to 85.7% on others. Despite these inaccuracies, the models identified the nearest cardinal direction in most cases, suggesting associative learning, embodying human-like misconceptions. We discuss the potential of text-based data representing geographic relationships directly to improve the spatial reasoning capabilities of LLMs.
翻译:我们提出了一个评估大型语言模型(LLMs)判断地理位置间二级方向能力的基准,并将其应用于三个主流LLM:GPT-3.5、GPT-4和Llama-2。该基准专门评估LLM是否表现出类似人类的分层空间偏差,即对单个位置空间关系的判断会受到包含它们的更大群体之感知关系的影响。为探究此问题,我们设计了14个聚焦于美国知名城市的问题。其中7个问题旨在通过可能受更大地理单元(如州或国家)朝向影响的场景挑战LLM,其余7个问题则针对不易受此类层级分类影响的位置。在测试模型中,GPT-4表现最优,准确率达55.3%,其次是GPT-3.5(47.3%)和Llama-2(44.7%)。模型在疑似存在分层偏差的任务中准确率显著下降,例如GPT-4在此类任务上的准确率降至32.9%,而在其他任务上为85.7%。尽管存在这些不准确性,模型在大多数情况下仍能识别出最近的基本方向,这表明其具备联想学习能力,并体现出类似人类的认知误区。我们讨论了直接利用基于文本的地理关系数据来提升LLM空间推理能力的潜力。