Question-answering for domain-specific applications has recently attracted much interest due to the latest advancements in large language models (LLMs). However, accurately assessing the performance of these applications remains a challenge, mainly due to the lack of suitable benchmarks that effectively simulate real-world scenarios. To address this challenge, we introduce two product question-answering (QA) datasets focused on Adobe Acrobat and Photoshop products to help evaluate the performance of existing models on domain-specific product QA tasks. Additionally, we propose a novel knowledge-driven RAG-QA framework to enhance the performance of the models in the product QA task. Our experiments demonstrated that inducing domain knowledge through query reformulation allowed for increased retrieval and generative performance when compared to standard RAG-QA methods. This improvement, however, is slight, and thus illustrates the challenge posed by the datasets introduced.
翻译:针对特定领域应用的问答系统,由于大型语言模型(LLM)的最新进展,近来引起了广泛关注。然而,准确评估这些应用的性能仍然是一个挑战,这主要是由于缺乏能够有效模拟真实场景的合适基准。为应对这一挑战,我们引入了两个专注于 Adobe Acrobat 和 Photoshop 产品的问答(QA)数据集,以帮助评估现有模型在特定领域产品 QA 任务上的性能。此外,我们提出了一种新颖的知识驱动型 RAG-QA 框架,以提升模型在产品 QA 任务中的表现。我们的实验表明,与标准的 RAG-QA 方法相比,通过查询重构引入领域知识可以提高检索和生成性能。然而,这种改进是轻微的,从而说明了所引入数据集带来的挑战。