Large language models (LLM) have become state of the art in many benchmarks and conversational LLM applications like ChatGPT are now widely used by the public. Those LLMs can be used to generate large amounts of content which is posted on the internet to various platforms. As LLMs are trained on datasets usually collected from the internet, this LLM-generated content might be used to train the next generation of LLMs. Therefore, a self-consuming training loop emerges in which new LLM generations are trained on the output from the previous generations. We empirically study this self-consuming training loop using a novel dataset to analytically and accurately measure quality and diversity of generated outputs. We find that this self-consuming training loop initially improves both quality and diversity. However, after a few generations the output inevitably degenerates in diversity. We find that the rate of degeneration depends on the proportion of real and generated data.
翻译:大型语言模型(LLM)已在众多基准测试中达到最先进水平,且诸如ChatGPT等会话式LLM应用如今已被公众广泛使用。这些LLM可生成大量内容,并被发布到互联网上的各类平台。由于LLM通常以从互联网收集的数据集进行训练,这些由LLM生成的内容可能用于训练下一代LLM。因此,一种自我消耗训练循环随之出现:新一代LLM基于前一代的输出进行训练。我们通过一个新颖的数据集对该自我消耗训练循环进行实证研究,以分析并精确衡量生成输出的质量和多样性。我们发现,这种自我消耗训练循环最初会同时提升质量和多样性。然而,经过几代之后,输出必然在多样性上退化。我们还发现,退化速率取决于真实数据与生成数据的比例。