Most Machine Learning research evaluates the best solutions in terms of performance. However, in the race for the best performing model, many important aspects are often overlooked when, on the contrary, they should be carefully considered. In fact, sometimes the gaps in performance between different approaches are neglectable, whereas factors such as production costs, energy consumption, and carbon footprint must take into consideration. Large Language Models (LLMs) are extensively adopted to address NLP problems in academia and industry. In this work, we present a detailed quantitative comparison of LLM and traditional approaches (e.g. SVM) on the LexGLUE benchmark, which takes into account both performance (standard indices) and alternative metrics such as timing, power consumption and cost, in a word: the carbon-footprint. In our analysis, we considered the prototyping phase (model selection by training-validation-test iterations) and in-production phases separately, since they follow different implementation procedures and also require different resources. The results indicate that very often, the simplest algorithms achieve performance very close to that of large LLMs but with very low power consumption and lower resource demands. The results obtained could suggest companies to include additional evaluations in the choice of Machine Learning (ML) solutions.
翻译:大多数机器学习研究以性能为标准评估最佳解决方案。然而,在追求最优模型的竞赛中,许多重要方面常被忽视,而这些因素本应得到审慎考量。事实上,不同方法之间的性能差异有时微不足道,而生产成本、能耗和碳足迹等要素却必须纳入考虑。大型语言模型(LLMs)已被广泛用于解决学术界和工业界的自然语言处理问题。本研究基于LexGLUE基准,对LLMs与传统方法(如SVM)进行了详细的定量比较,该比较同时考虑了性能(标准指标)和替代性度量指标(如时间、功耗和成本),概括而言即碳足迹。在分析中,我们分别考虑了原型设计阶段(通过训练-验证-测试迭代进行模型选择)和生产阶段,因为二者遵循不同的实施流程且所需资源各异。结果表明,最简单的算法通常能以极低的功耗和资源需求达到接近大型LLMs的性能。这一发现可能促使企业将更多评估维度纳入机器学习解决方案的选择过程中。