The advancement of large language models (LLMs) has enhanced tabular question answering (Tabular QA), yet they struggle with open-domain queries exhibiting underspecified or uncertain expressions. To address this, we introduce the ODUTQA-MDC task and the first comprehensive benchmark to tackle it. This benchmark includes: (1) a large-scale ODUTQA dataset with 209 tables and 25,105 QA pairs; (2) a fine-grained labeling scheme for detailed evaluation; and (3) a dynamic clarification interface that simulates user feedback for interactive assessment. We also propose MAIC-TQA, a multi-agent framework that excels at detecting ambiguities, clarifying them through dialogue, and refining answers. Experiments validate our benchmark and framework, establishing them as a key resource for advancing conversational, underspecification-aware Tabular QA research.
翻译:大型语言模型的进步增强了表格问答能力,但在处理存在未完全指定或不确定表达的开放域查询时仍面临挑战。为此,我们引入ODUTQA-MDC任务及其首个综合性基准以应对该问题。该基准包含:(1)含209张表格与25,105个问答对的大规模ODUTQA数据集;(2)用于细粒度评估的精细化标注方案;(3)模拟用户反馈的动态澄清接口以实现交互式测评。我们同时提出MAIC-TQA多智能体框架,该框架擅长检测歧义、通过对话消除歧义并优化答案。实验验证了我们的基准与框架,使其成为推动对话式、感知未完全指定特征的表格问答研究的关键资源。