The widespread adoption of Large Language Models and publicly available ChatGPT has marked a significant turning point in the integration of Artificial Intelligence into people's everyday lives. The academic community has taken notice of these technological advancements and has expressed concerns regarding the difficulty of discriminating between what is real and what is artificially generated. Thus, researchers have been working on developing effective systems to identify machine-generated text. In this study, we utilize the GPT-3 model to generate scientific paper abstracts through Artificial Intelligence and explore various text representation methods when combined with Machine Learning models with the aim of identifying machine-written text. We analyze the models' performance and address several research questions that rise during the analysis of the results. By conducting this research, we shed light on the capabilities and limitations of Artificial Intelligence generated text.
翻译:大语言模型的广泛采用以及公开可用的ChatGPT标志着人工智能融入人们日常生活的重大转折点。学术界已注意到这些技术进步,并对区分真实内容与人工生成内容的难度表示担忧。因此,研究人员一直致力于开发有效的系统以识别机器生成的文本。在本研究中,我们利用GPT-3模型通过人工智能生成科学论文摘要,并探索结合机器学习模型时的各种文本表示方法,旨在识别机器撰写的文本。我们分析了模型的性能,并探讨了在结果分析过程中出现的若干研究问题。通过开展此项研究,我们揭示了人工智能生成文本的能力与局限性。