Rapid advancements over the years have helped machine learning models reach previously hard-to-achieve goals, sometimes even exceeding human capabilities. However, to attain the desired accuracy, the model sizes and in turn their computational requirements have increased drastically. Thus, serving predictions from these models to meet any target latency and cost requirements of applications remains a key challenge, despite recent work in building inference-serving systems as well as algorithmic approaches that dynamically adapt models based on inputs. In this paper, we introduce a form of dynamism, modality selection, where we adaptively choose modalities from inference inputs while maintaining the model quality. We introduce MOSEL, an automated inference serving system for multi-modal ML models that carefully picks input modalities per request based on user-defined performance and accuracy requirements. MOSEL exploits modality configurations extensively, improving system throughput by 3.6$\times$ with an accuracy guarantee and shortening job completion times by 11$\times$.
翻译:近年来,机器学习模型的快速进步使其达到了以往难以企及的目标,有时甚至超越了人类能力。然而,为获得期望精度,模型规模及其计算需求急剧增长。因此,尽管在构建推理服务系统及基于输入动态调整模型的算法方法方面已有最新研究,如何从这些模型提供满足应用延迟与成本目标的预测服务仍是一项关键挑战。本文引入一种新的动态机制——模态选择,通过自适应地从推理输入中选取模态,同时保持模型质量。我们提出MOSEL,一种面向多模态机器学习模型的自动化推理服务系统,其能根据用户定义的性能与精度要求,为每个请求精细选择输入模态。MOSEL充分挖掘模态配置潜力,在保证精度前提下将系统吞吐量提升3.6倍,并将作业完成时间缩短11倍。