LAraBench: Benchmarking Arabic AI with Large Language Models

Ahmed Abdelali,Hamdy Mubarak,Shammur Absar Chowdhury,Maram Hasanain,Basel Mousi,Sabri Boughorbel,Yassine El Kheir,Daniel Izham,Fahim Dalvi,Majd Hawasly,Nizi Nazar,Yousseif Elshahawy,Ahmed Ali,Nadir Durrani,Natasa Milic-Frayling,Firoj Alam

from arxiv, Foundation Models, Large Language Models, Arabic NLP, Arabic Speech, Arabic AI, GPT3.5 Evaluation, USM Evaluation, Whisper Evaluation, GPT-4, BLOOMZ, Jais13b

Recent advancements in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. Despite this progress, these models lack specific benchmarking against state-of-the-art (SOTA) models tailored to particular languages and tasks. LAraBench addresses this gap for Arabic Natural Language Processing (NLP) and Speech Processing tasks, including sequence tagging and content classification across different domains. We utilized models such as GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM, employing zero and few-shot learning techniques to tackle 33 distinct tasks across 61 publicly available datasets. This involved 98 experimental setups, encompassing ~296K data points, ~46 hours of speech, and 30 sentences for Text-to-Speech (TTS). This effort resulted in 330+ sets of experiments. Our analysis focused on measuring the performance gap between SOTA models and LLMs. The overarching trend observed was that SOTA models generally outperformed LLMs in zero-shot learning, with a few exceptions. Notably, larger computational models with few-shot learning techniques managed to reduce these performance gaps. Our findings provide valuable insights into the applicability of LLMs for Arabic NLP and speech processing tasks.

翻译：近年来，大型语言模型（LLM）的进展显著影响了语言与语音研究领域。然而，这些模型缺乏针对特定语言和任务的、与当前最优（SOTA）模型进行对比的专项基准测试。LAraBench填补了这一空白，聚焦阿拉伯语自然语言处理（NLP）和语音处理任务，涵盖不同领域的序列标注与内容分类。我们采用GPT-3.5-turbo、GPT-4、BLOOMZ、Jais-13b-chat、Whisper和USM等模型，结合零样本与少样本学习技术，在61个公开数据集上处理了33项不同任务。实验设置共98组，涉及约29.6万个数据点、约46小时语音数据及30句文本转语音（TTS）样本，最终完成330余组实验。我们重点分析了SOTA模型与LLM之间的性能差距。总体趋势表明，SOTA模型在零样本学习中通常优于LLM，仅存在少数例外。值得注意的是，采用少样本学习技术的大规模计算模型能够缩小这一性能差距。本研究的发现为LLM在阿拉伯语NLP及语音处理任务中的适用性提供了重要洞见。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

O’Reilly报告：知识图谱崛起——面向现代数据集成和数据结构体系，“The Rise of the Knowledge Graph——Toward Modern Data Integration and the Data Fabric Architecture”

专知会员服务

49+阅读 · 2022年2月18日

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

32+阅读 · 2021年9月29日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日