Italy is characterized by a one-of-a-kind linguistic diversity landscape in Europe, which implicitly encodes local knowledge, cultural traditions, artistic expressions and history of its speakers. However, most local languages and dialects in Italy are at risk of disappearing within few generations. The NLP community has recently begun to engage with endangered languages, including those of Italy. Yet, most efforts assume that these varieties are under-resourced language monoliths with an established written form and homogeneous functions and needs, and thus highly interchangeable with each other and with high-resource, standardized languages. In this paper, we introduce the linguistic context of Italy and challenge the default machine-centric assumptions of NLP for Italy's language varieties. We advocate for a shift in the paradigm from machine-centric to speaker-centric NLP, and provide recommendations and opportunities for work that prioritizes languages and their speakers over technological advances. To facilitate the process, we finally propose building a local community towards responsible, participatory efforts aimed at supporting vitality of languages and dialects of Italy.
翻译:意大利在欧洲拥有独一无二的语言多样性景观,其语言隐含着当地知识、文化传统、艺术表达及使用者的历史。然而,意大利的大多数地方语言和方言面临在几代人内消失的风险。自然语言处理(NLP)社区近期开始关注濒危语言,包括意大利的语言变体。但大多数研究将这些变体视为缺乏资源的语言单体,假定它们具备固定的书面形式、同质的功能与需求,因此彼此之间以及与高资源标准化语言之间高度可互换。本文引入意大利的语言背景,质疑NLP对意大利语言变体的默认以机器为中心的假设。我们倡导从以机器为中心转向以说话者为中心的NLP范式,并提供优先考虑语言及其使用者(而非技术进展)的建议与机遇。为推动这一进程,我们最终提议建立本地社区,开展负责任的、参与性的工作,以支持意大利语言和方言的活力。