Compound AI systems (CASs) that employ LLMs as agents to accomplish knowledge-intensive tasks via interactions with tools and data retrievers have garnered significant interest within database and AI communities. While these systems have the potential to supplement typical analysis workflows of data analysts in enterprise data platforms, unfortunately, CASs are subject to the same data discovery challenges that analysts have encountered over the years -- silos of multimodal data sources, created across teams and departments within an organization, make it difficult to identify appropriate data sources for accomplishing the task at hand. Existing data discovery benchmarks do not model such multimodality and multiplicity of data sources. Moreover, benchmarks of CASs prioritize only evaluating end-to-end task performance. To catalyze research on evaluating the data discovery performance of multimodal data retrievers in CASs within a real-world setting, we propose CMDBench, a benchmark modeling the complexity of enterprise data platforms. We adapt existing datasets and benchmarks in open-domain -- from question answering and complex reasoning tasks to natural language querying over structured data -- to evaluate coarse- and fine-grained data discovery and task execution performance. Our experiments reveal the impact of data retriever design on downstream task performance -- a 46% drop in task accuracy on average -- across various modalities, data sources, and task difficulty. The results indicate the need to develop optimization strategies to identify appropriate LLM agents and retrievers for efficient execution of CASs over enterprise data.
翻译:采用大型语言模型(LLM)作为智能体、通过工具与数据检索器交互以完成知识密集型任务的复合AI系统(CAS),已在数据库与人工智能领域引起广泛关注。尽管这类系统有望增强企业数据平台中数据分析师的典型工作流程,但不幸的是,CAS同样面临着分析师多年来遭遇的数据发现挑战——组织内跨团队与部门创建的多模态数据源孤岛,使得难以识别适用于当前任务的恰当数据源。现有数据发现基准测试未能建模此类多模态与多源数据特性。此外,现有CAS基准测试仅侧重于端到端任务性能评估。为促进在真实场景下评估CAS中多模态数据检索器数据发现性能的研究,我们提出了CMDBench——一个模拟企业数据平台复杂性的基准测试框架。我们通过改造开放域现有数据集与基准测试(涵盖问答与复杂推理任务至结构化数据的自然语言查询),以评估粗粒度与细粒度的数据发现及任务执行性能。实验结果表明,在不同模态、数据源及任务难度下,数据检索器设计对下游任务性能存在显著影响——平均任务准确率下降达46%。该结果揭示了需要开发优化策略,以识别合适的LLM智能体与检索器,从而在企业数据上高效执行CAS。