Many recent improvements in NLP stem from the development and use of large pre-trained language models (PLMs) with billions of parameters. Large model sizes makes computational cost one of the main limiting factors for training and evaluating such models; and has raised severe concerns about the sustainability, reproducibility, and inclusiveness for researching PLMs. These concerns are often based on personal experiences and observations. However, there had not been any large-scale surveys that investigate them. In this work, we provide a first attempt to quantify these concerns regarding three topics, namely, environmental impact, equity, and impact on peer reviewing. By conducting a survey with 312 participants from the NLP community, we capture existing (dis)parities between different and within groups with respect to seniority, academia, and industry; and their impact on the peer reviewing process. For each topic, we provide an analysis and devise recommendations to mitigate found disparities, some of which already successfully implemented. Finally, we discuss additional concerns raised by many participants in free-text responses.
翻译:自然语言处理领域的许多最新进展源于对具有数十亿参数的大型预训练语言模型(PLMs)的开发与应用。巨大的模型规模使得计算成本成为训练和评估这类模型的主要限制因素之一,并引发了对PLM研究可持续性、可重复性和包容性的严重担忧。这些担忧往往基于个人经验与观察,但此前从未有过大规模调查研究对其加以验证。本文首次尝试量化这些担忧,具体围绕三个主题展开:环境影响、公平性以及对同行评议的影响。通过对NLP领域312名参与者进行问卷调查,我们捕捉了不同群体内部及之间在资历水平、学术界与工业界背景等方面的现有(不)平等现象,以及这些现象对同行评审过程的影响。针对每个主题,我们进行了分析并提出缓解已发现不平等的建议,部分建议已成功实施。最后,我们讨论了参与者在自由文本回复中提出的其他担忧。