When Automated Assessment Meets Automated Content Generation: Examining Text Quality in the Era of GPTs

The use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility assessment of online content. A significant disruption at the intersection of ML and text are text-generating large-language models such as generative pre-trained transformers (GPTs). We empirically assess the differences in how ML-based scoring models trained on human content assess the quality of content generated by humans versus GPTs. To do so, we propose an analysis framework that encompasses essay scoring ML-models, human and ML-generated essays, and a statistical model that parsimoniously considers the impact of type of respondent, prompt genre, and the ML model used for assessment model. A rich testbed is utilized that encompasses 18,460 human-generated and GPT-based essays. Results of our benchmark analysis reveal that transformer pretrained language models (PLMs) more accurately score human essay quality as compared to CNN/RNN and feature-based ML methods. Interestingly, we find that the transformer PLMs tend to score GPT-generated text 10-15\% higher on average, relative to human-authored documents. Conversely, traditional deep learning and feature-based ML models score human text considerably higher. Further analysis reveals that although the transformer PLMs are exclusively fine-tuned on human text, they more prominently attend to certain tokens appearing only in GPT-generated text, possibly due to familiarity/overlap in pre-training. Our framework and results have implications for text classification settings where automated scoring of text is likely to be disrupted by generative AI.

翻译：使用机器学习（ML）模型评估和评分文本数据在自然语言处理、信息检索、搜索与推荐以及在线内容可信度评估等多个领域日益普及。ML与文本交叉领域的一项重要变革是文本生成式大型语言模型，例如生成式预训练转换器（GPT）。本文实证评估了基于人类内容训练的ML评分模型在评估人类与GPT生成内容质量方面的差异。为此，我们提出了一种分析框架，该框架整合了论文评分ML模型、人类与ML生成的论文，以及一个简约考量受访者类型、提示体裁和评估模型所用ML模型影响的统计模型。我们利用了一个包含18,460篇人类与GPT生成论文的丰富测试平台。基准分析结果表明，与CNN/RNN及基于特征的ML方法相比，转换器预训练语言模型（PLM）能更准确地评估人类论文质量。有趣的是，我们发现转换器PLM对GPT生成文本的评分平均比人类撰写文本高出10-15%。相反，传统深度学习与基于特征的ML模型则对人类文本的评分显著更高。进一步分析揭示，尽管转换器PLM仅针对人类文本进行微调，但它们更显著地关注仅出现在GPT生成文本中的特定标记，这可能源于预训练中的熟悉度/重叠。我们的框架与结果对文本分类场景具有启示意义，在这些场景中，文本的自动化评估很可能受到生成式AI的冲击。

相关内容

Automator

关注 5

Automator是苹果公司为他们的Mac OS X系统开发的一款软件。 只要通过点击拖拽鼠标等操作就可以将一系列动作组合成一个工作流，从而帮助你自动的（可重复的）完成一些复杂的工作。Automator还能横跨很多不同种类的程序，包括：查找器、Safari网络浏览器、iCal、地址簿或者其他的一些程序。它还能和一些第三方的程序一起工作，如微软的Office、Adobe公司的Photoshop或者Pixelmator等。

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

32+阅读 · 2021年9月29日

【亚马逊-WWW2020】不解析,生成!用于面向任务的语义分析的序列到序列体系结构，Don't Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing

专知会员服务

15+阅读 · 2020年2月1日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日