The proliferation of open large language models (LLMs) is fostering a vibrant ecosystem in artificial intelligence (AI). However, the methods of collaboration used to develop open LLMs, both before and after their public release, have not yet been systematically studied, limiting our understanding of how open LLM projects are initiated, organised, and governed, as well as the opportunities to further foster this ecosystem. We address this gap through an exploratory analysis of open collaboration throughout the development and reuse lifecycle of open LLMs, drawing on semi-structured interviews with the developers of 14 diverse open LLM projects. These collaborations span multiple artefact domains -- including models, data, software, evaluation, compute, and community engagement -- each enabling distinct forms of participation and involving different stakeholders that evolves across the LLM development lifecycle, shifting from concentrated, selective engagement in the early stages to broader, distributed participation after model release. The open LLM developers are motivated by a variety of social, economic, and technological motivations, ranging from democratising access to AI and promoting open science to building regional ecosystems and expanding language representation. These dynamics are coordinated through a range of governance structures, typically formal and professionalised to varying degrees, including centralised company-led efforts to decentralised grassroots initiatives. We synthesise our findings in a conceptual model of open collaboration in open LLM ecosystems, provide recommendations for practice, and conclude that openness in open source AI is not a uniform property but an emergent outcome of how collaboration is organised across interconnected artefact domains, lifecycle stages, and institutional contexts.
翻译:开放大语言模型(LLMs)的涌现正培育着人工智能(AI)领域充满活力的生态系统。然而,用于开发开放大语言模型的协作方式——无论是在其公开发布之前还是之后——尚未得到系统性研究,这限制了我们对于开放大语言模型项目如何启动、组织与治理的理解,也制约了进一步培育该生态系统的机遇。我们通过一项探索性分析,聚焦于开放大语言模型开发与复用全生命周期中的开放协作,以填补这一空白。该分析基于对14个多元化开放大语言模型项目开发者进行的半结构化访谈。这些协作跨越多个制品领域——包括模型、数据、软件、评估、计算和社区参与——每个领域都支持不同的参与形式,并涉及随时间演变的利益相关方,其动态从早期的集中化、选择性参与,转向模型发布后更广泛、分布式的参与。开放大语言模型开发者受到社会、经济及技术等多重动机的驱动,涵盖从普及AI接入、推动开放科学,到构建区域生态系统与扩展语言表征等目标。这些动态通过一系列正式化与专业化程度各异的治理结构进行协调,包括集中化的公司主导型举措与分散化的草根型倡议。我们将研究发现综合成一个关于开放大语言模型生态系统中开放协作的概念模型,并为实践提供建议,进而指出:在开源AI中,开放性并非一个统一属性,而是协作如何跨相互关联的制品领域、生命周期阶段及制度背景进行组织而产生的新兴结果。