Large language models (LLMs) with in-context learning have demonstrated impressive generalization capabilities in the cross-domain text-to-SQL task, without the use of in-domain annotations. However, incorporating in-domain demonstration examples has been found to greatly enhance LLMs' performance. In this paper, we delve into the key factors within in-domain examples that contribute to the improvement and explore whether we can harness these benefits without relying on in-domain annotations. Based on our findings, we propose a demonstration selection framework ODIS which utilizes both out-of-domain examples and synthetically generated in-domain examples to construct demonstrations. By retrieving demonstrations from hybrid sources, ODIS leverages the advantages of both, showcasing its effectiveness compared to baseline methods that rely on a single data source. Furthermore, ODIS outperforms state-of-the-art approaches on two cross-domain text-to-SQL datasets, with improvements of 1.1 and 11.8 points in execution accuracy, respectively.
翻译:大型语言模型(LLMs)通过上下文学习在跨领域文本到SQL任务中展现出了令人印象深刻的泛化能力,且无需使用领域内标注。然而,研究表明融入领域内示范示例能显著提升LLMs的性能。本文深入探究领域内示例中促进性能提升的关键因素,并探索是否能在不依赖领域内标注的情况下利用这些优势。基于研究发现,我们提出示范选择框架ODIS,该框架综合使用领域外示例与合成生成的领域内示例来构建示范。通过从混合数据源中检索示范,ODIS融合了两者的优势,相较于依赖单一数据源的基线方法展现了更优效果。此外,ODIS在两个跨领域文本到SQL数据集上分别以1.1和11.8个百分点的执行准确率提升超越了现有最优方法。