The advancement of large language models (LLMs) has enhanced tabular question answering (Tabular QA), yet they struggle with open-domain queries exhibiting underspecified or uncertain expressions. To address this, we introduce the ODUTQA-MDC task and the first comprehensive benchmark to tackle it. This benchmark includes: (1) a large-scale ODUTQA dataset with 209 tables and 25,105 QA pairs; (2) a fine-grained labeling scheme for detailed evaluation; and (3) a dynamic clarification interface that simulates user feedback for interactive assessment. We also propose MAIC-TQA, a multi-agent framework that excels at detecting ambiguities, clarifying them through dialogue, and refining answers. Experiments validate our benchmark and framework, establishing them as a key resource for advancing conversational, underspecification-aware Tabular QA research.
翻译:大型语言模型(LLM)的进步提升了表格问答(Tabular QA)的性能,但它们在处理包含未明确指定或不确定表述的开放域查询时仍存在困难。为解决这一问题,我们提出了ODUTQA-MDC任务及其首个综合基准测试。该基准测试包括:(1)包含209个表格和25,105个问答对的大规模ODUTQA数据集;(2)用于细粒度评估的精细标注方案;(3)可模拟用户反馈进行交互评估的动态澄清界面。我们还提出了MAIC-TQA——一种能够有效检测歧义、通过对话澄清歧义并优化答案的多智能体框架。实验验证了我们的基准测试和框架,使其成为推动具歧义感知能力的对话式表格问答研究的关键资源。