基于大语言模型的Python库迁移实证研究 (An Empirical Study of Python Library Migration Using Large Language Models)

Library migration is the process of replacing one library with another library that provides similar functionality. Manual library migration is time consuming and error prone, as it requires developers to understand the APIs of both libraries, map them, and perform the necessary code transformations. Large Language Models (LLMs) are shown to be effective at generating and transforming code as well as finding similar code, which are necessary upstream tasks for library migration. Such capabilities suggest that LLMs may be suitable for library migration. Accordingly, this paper investigates the effectiveness of LLMs for migration between Python libraries. We evaluate three LLMs, Llama 3.1, GPT-4o mini, and GPT-4o on PyMigBench, where we migrate 321 real-world library migrations that include 2,989 migration-related code changes. To measure correctness, we (1) compare the LLM's migrated code with the developers' migrated code in the benchmark and (2) run the unit tests available in the client repositories. We find that LLama 3.1, GPT-4o mini, and GPT-4o correctly migrate 89%, 89%, and 94% of the migration-related code changes, respectively. We also find that 36%, 52% and 64% of the LLama 3.1, GPT-4o mini, and GPT-4o migrations pass the same tests that passed in the developer's migration. To ensure the LLMs are not reciting the migrations, we also evaluate them on 10 new repositories where the migration never happened. Overall, our results suggest that LLMs can be effective in migrating code between libraries, but we also identify some open challenges.

翻译：库迁移是指将一个库替换为另一个提供相似功能的库的过程。手动库迁移耗时且容易出错，因为它要求开发者理解两个库的API、建立映射关系并执行必要的代码转换。大语言模型（LLMs）已被证明在生成和转换代码以及查找相似代码方面表现有效，而这些正是库迁移所需的上游任务。此类能力表明LLMs可能适用于库迁移。为此，本文研究了LLMs在Python库间迁移的有效性。我们在PyMigBench基准上评估了三个LLMs：Llama 3.1、GPT-4o mini和GPT-4o，其中我们迁移了321个真实世界的库迁移案例，包含2,989个与迁移相关的代码变更。为衡量正确性，我们（1）将LLM迁移后的代码与基准中开发者迁移的代码进行比较，以及（2）运行客户端仓库中可用的单元测试。我们发现Llama 3.1、GPT-4o mini和GPT-4o分别正确迁移了89%、89%和94%的迁移相关代码变更。我们还发现，由Llama 3.1、GPT-4o mini和GPT-4o完成的迁移中，分别有36%、52%和64%通过了与开发者迁移相同的测试。为确保LLMs并非在复述已有的迁移，我们还在10个从未发生过迁移的新仓库上对它们进行了评估。总体而言，我们的结果表明LLMs在库间代码迁移方面可以发挥有效作用，但我们也发现了一些尚待解决的挑战。