Deploying Large Language Model (LLM) applications, particularly those relying on Retrieval-Augmented Generation (RAG), remains challenging due to high computational demands, outdated knowledge bases, and the need to manually select optimal pipeline components. In this work, we propose a modular framework for benchmarking and guiding the efficient development of RAG applications by focusing on resource telemetry and component recommendation, suggesting the best components for a domain-specific dataset. Our approach leverages core techniques in LLM applications, including document chunking, vector databases, embedding models, and retrievers, to evaluate trade-offs among accuracy, efficiency, and scalability. By directly correlating retrieval and generation quality with underlying hardware constraints, RAGe supports researchers to identify the most effective, domain-specific RAG setups for their specific operational needs, facilitating rapid prototyping even on consumer-grade hardware.
翻译:部署大语言模型(LLM)应用,尤其是依赖检索增强生成(RAG)的应用,仍面临挑战,原因包括高计算需求、知识库过时以及需要手动选择最优流水线组件。本文提出了一种模块化框架,通过聚焦资源遥测与组件推荐,为领域特定数据集推荐最佳组件,从而对标并引导RAG应用的高效开发。我们的方法利用LLM应用中的核心技术(包括文档分块、向量数据库、嵌入模型和检索器),评估准确性、效率与可扩展性之间的权衡。通过直接关联检索与生成质量与底层硬件约束,RAGe能够帮助研究人员为其特定操作需求识别最有效的、领域特定的RAG配置,即使在消费级硬件上也能促进快速原型开发。