In this paper, we investigate the intersection of large generative AI models and cloud-native computing architectures. Recent large models such as ChatGPT, while revolutionary in their capabilities, face challenges like escalating costs and demand for high-end GPUs. Drawing analogies between large-model-as-a-service (LMaaS) and cloud database-as-a-service (DBaaS), we describe an AI-native computing paradigm that harnesses the power of both cloud-native technologies (e.g., multi-tenancy and serverless computing) and advanced machine learning runtime (e.g., batched LoRA inference). These joint efforts aim to optimize costs-of-goods-sold (COGS) and improve resource accessibility. The journey of merging these two domains is just at the beginning and we hope to stimulate future research and development in this area.
翻译:本文探讨了大生成式AI模型与云原生计算架构的交汇点。近期如ChatGPT等大模型虽在能力上具有革命性,但面临成本攀升与高端GPU需求等挑战。通过类比大模型即服务(LMaaS)与云数据库即服务(DBaaS),我们描述了一种融合云原生技术(如多租户与无服务器计算)与先进机器学习运行时(如批量化LoRA推理)的AI原生计算范式。这些协同努力旨在优化商品销售成本(COGS)并提升资源可及性。将这两个领域融合的探索尚处于起步阶段,我们期望能激发该领域的未来研究与发展。