We demonstrate LLARS (LLM Assisted Research System), an open-source platform that bridges the gap between domain experts and developers for building LLM-based systems. It integrates three tightly connected modules into an end-to-end pipeline: Collaborative Prompt Engineering for real-time co-authoring with version control and instant LLM testing, Batch Generation for configurable output production across user-selected prompts $\times$ models $\times$ data with cost control, and Hybrid Evaluation where human and LLM evaluators jointly assess outputs through diverse assessment methods, with live agreement metrics and provenance analysis to identify the best model-prompt combination for a given use case. New prompts and models are automatically available for batch generation and completed batches can be turned into evaluation scenarios with a single click. Interviews with six domain experts and three developers in online counselling confirmed that LLARS feels intuitive, saves considerable time by keeping everything in one place and makes interdisciplinary collaboration seamless.
翻译:我们展示了LLARS(大语言模型辅助研究系统),这是一个开源平台,旨在弥合领域专家与开发者之间在构建基于大语言模型的系统方面的鸿沟。该系统将三个紧密相连的模块整合为一个端到端流水线:协作式提示工程,支持带版本控制的实时协同编写及即时大语言模型测试;批量生成,可在用户选定提示×模型×数据组合上配置输出生成并控制成本;混合评估,通过多样化评估方法由人类与LLM评估员联合评估输出,并提供实时代码一致性指标与来源分析,以识别特定用例的最佳模型-提示组合。新提示与模型可自动用于批量生成,已完成批处理可一键转化为评估场景。对在线咨询领域的六位领域专家和三位开发者的访谈证实,LLARS操作直观、通过将所有功能集中一处显著节省时间,并使跨学科协作无缝衔接。