Natural Language Processing (NLP) technologies have revolutionized the way we interact with information systems, with a significant focus on converting natural language queries into formal query languages such as SQL. However, less emphasis has been placed on the Corpus Query Language (CQL), a critical tool for linguistic research and detailed analysis within text corpora. The manual construction of CQL queries is a complex and time-intensive task that requires a great deal of expertise, which presents a notable challenge for both researchers and practitioners. This paper presents the first text-to-CQL task that aims to automate the translation of natural language into CQL. We present a comprehensive framework for this task, including a specifically curated large-scale dataset and methodologies leveraging large language models (LLMs) for effective text-to-CQL task. In addition, we established advanced evaluation metrics to assess the syntactic and semantic accuracy of the generated queries. We created innovative LLM-based conversion approaches and detailed experiments. The results demonstrate the efficacy of our methods and provide insights into the complexities of text-to-CQL task.
翻译:自然语言处理技术已经革新了我们与信息系统交互的方式,其中大量研究聚焦于将自然语言查询转换为SQL等形式化查询语言。然而,鲜有关注语料库查询语言(CQL)——这一对语言学研究和文本语料库精细分析至关重要的工具。手动构建CQL查询是一项复杂且耗时的任务,需要大量专业知识,这对研究人员和从业者构成了显著挑战。本文首次提出文本到CQL任务,旨在自动化实现自然语言到CQL的转换。我们为该任务构建了一个完整框架,包括专门策划的大规模数据集,以及利用大语言模型(LLMs)实现高效文本到CQL转换的方法论。此外,我们建立了先进的评估指标,用于衡量生成查询的句法与语义准确性。我们创新性地提出了基于LLM的转换方法,并进行了详尽的实验。实验结果验证了方法的有效性,并揭示了文本到CQL任务的复杂性。