Recent advances in deep learning methods for natural language processing (NLP) have created new business opportunities and made NLP research critical for industry development. As one of the big players in the field of NLP, together with governments and universities, it is important to track the influence of industry on research. In this study, we seek to quantify and characterize industry presence in the NLP community over time. Using a corpus with comprehensive metadata of 78,187 NLP publications and 701 resumes of NLP publication authors, we explore the industry presence in the field since the early 90s. We find that industry presence among NLP authors has been steady before a steep increase over the past five years (180% growth from 2017 to 2022). A few companies account for most of the publications and provide funding to academic researchers through grants and internships. Our study shows that the presence and impact of the industry on natural language processing research are significant and fast-growing. This work calls for increased transparency of industry influence in the field.
翻译:近年来,深度学习在自然语言处理(NLP)方法上的进步创造了新的商业机会,并使NLP研究对产业发展至关重要。作为NLP领域的主要参与者之一,与政府和大学一样,追踪产业对研究的影响具有重要意义。本研究旨在量化和描述NLP社区中产业存在随时间的变化。通过使用包含78,187篇NLP出版物全面元数据的语料库以及701份NLP出版物作者的简历,我们探索了自90年代初以来该领域的产业存在情况。我们发现,NLP作者中的产业存在在过去五年急剧增长之前一直保持稳定(2017年至2022年间增长了180%)。少数公司贡献了大部分出版物,并通过资助和实习机会为学术研究人员提供资金支持。我们的研究表明,产业对自然语言处理研究的存在和影响是显著且快速增长的。这项工作呼吁提高该领域产业影响的透明度。