Text-to-SQL allows experts to use databases without in-depth knowledge of them. However, real-world tasks have both query and data ambiguities. Most works on Text-to-SQL focused on query ambiguities and designed chat interfaces for experts to provide clarifications. In contrast, the data management community has long studied data ambiguities, but mainly addresses error detection and correction, rather than documenting them for disambiguation in data tasks. This work delves into these data ambiguities in real-world datasets. We have identified prevalent data ambiguities of value consistency, data coverage, and data granularity that affect tasks. We examine how documentation, originally made to help humans to disambiguate data, can help GPT-4 with Text-to-SQL tasks. By offering documentation on these, we found GPT-4's performance improved by 28.9%.
翻译:文本到SQL技术使专家无需深入了解数据库即可使用数据库。然而,实际任务中同时存在查询歧义和数据歧义。大多数文本到SQL研究聚焦于查询歧义,并设计聊天界面供专家提供澄清说明。相比之下,数据管理领域长期研究数据歧义,但主要关注错误检测与修正,而非通过文档化来消除数据任务中的歧义。本研究深入探讨真实世界数据集中的数据歧义,识别出影响任务的价值一致性、数据覆盖范围与数据粒度三类常见数据歧义。我们考察了原本为帮助人类消除数据歧义而设计的文档化,如何辅助GPT-4完成文本到SQL任务。通过提供相关文档,GPT-4的性能提升了28.9%。