Large language models (LLMs) have been touted to enable increased productivity in many areas of today's work life. Scientific research as an area of work is no exception: the potential of LLM-based tools to assist in the daily work of scientists has become a highly discussed topic across disciplines. However, we are only at the very onset of this subject of study. It is still unclear how the potential of LLMs will materialise in research practice. With this study, we give first empirical evidence on the use of LLMs in the research process. We have investigated a set of use cases for LLM-based tools in scientific research, and conducted a first study to assess to which degree current tools are helpful. In this paper we report specifically on use cases related to software engineering, such as generating application code and developing scripts for data analytics. While we studied seemingly simple use cases, results across tools differ significantly. Our results highlight the promise of LLM-based tools in general, yet we also observe various issues, particularly regarding the integrity of the output these tools provide.
翻译:大语言模型(LLMs)已被认为能够提升当今工作生活中多个领域的效率。科学研究作为工作领域之一也不例外:基于LLM的工具在协助科学家日常工作中的潜力已成为跨学科高度讨论的话题。然而,我们仍处于该研究主题的起步阶段。LLMs的潜力将如何在研究实践中具体化尚不明确。通过本研究,我们首次提供了关于LLMs在研究过程中使用的实证证据。我们研究了基于LLM的工具在科学研究中的一组应用案例,并开展了初步研究以评估当前工具的实用程度。本文具体报告了与软件工程相关的应用案例,例如生成应用程序代码和开发数据分析脚本。尽管我们研究了看似简单的案例,不同工具的结果存在显著差异。研究结果凸显了基于LLM工具的总体前景,但我们也观察到各种问题,尤其是这些工具所提供输出的完整性方面。