Multi-party dialogue is a critical setting for studying collaborative reasoning and decision-making, yet existing datasets rarely focus on structured, in-depth complex reasoning tasks. We introduce DeliChess, a novel dataset of group deliberation dialogues in which participants collaboratively solve multiple-choice chess puzzles. Each group first completes the puzzle individually, then engages in a multi-party discussion before submitting a revised collective answer. The dataset includes 107 dialogues with full transcripts, pre- and post-discussion choices, and metadata on puzzle difficulty and move quality. We evaluate performance using three metrics based on chess engine evaluations, and find that deliberation significantly improves group accuracy. We further analyse the role of probing utterances (i.e., messages that elicit proposals, justifications, or strategic reflection) using a classifier trained on prior deliberation data. While probing makes group performance more variable after discussion, it does not consistently lead to better performance. Our dataset offers a rich testbed for modelling group reasoning, dialogue dynamics, and the resolution of differing perspectives and opinions in a well-defined strategic domain.
翻译:多方对话是研究协作推理与决策制定的关键场景,然而现有数据集鲜少关注结构化、深层次的复杂推理任务。我们提出DeliChess——一个全新的群体深度讨论对话数据集,其中参与者通过协作方式解决多项选择国际象棋谜题。实验流程中,各小组先各自独立完成谜题,随后进行多方讨论,最终提交修正后的集体答案。该数据集包含107组对话文本、讨论前后选择记录,以及谜题难度与走棋质量的元数据。我们采用基于象棋引擎评估的三种指标进行性能评测,发现深度讨论显著提升了小组准确率。进一步地,我们利用先前深度讨论数据训练的预分类器,分析了试探性话语(即引发提案、论证或策略反思的发言)的作用。研究发现:试探性话语虽使讨论后的小组表现波动性增大,但并未持续带来性能提升。本数据集为在明确定义的策略领域中建模群体推理、对话动态以及不同观点分歧的消解机制提供了丰富的实验平台。