Large language models are widespread, with their performance on benchmarks frequently guiding user preferences for one model over another. However, the vast amount of data these models are trained on can inadvertently lead to contamination with public benchmarks, thus compromising performance measurements. While recently developed contamination detection methods try to address this issue, they overlook the possibility of deliberate contamination by malicious model providers aiming to evade detection. We argue that this setting is of crucial importance as it casts doubt on the reliability of public benchmarks. To more rigorously study this issue, we propose a categorization of both model providers and contamination detection methods. This reveals vulnerabilities in existing methods that we exploit with EAL, a simple yet effective contamination technique that significantly inflates benchmark performance while completely evading current detection methods.
翻译:大语言模型已广泛普及,其在基准测试上的表现常引导用户偏好选择某一模型。然而,这些模型训练时所依赖的海量数据可能不经意间混入公开基准测试的污染数据,从而损害性能测量的准确性。尽管近期开发的污染检测方法试图解决此问题,但它们忽略了模型提供者可能蓄意通过污染来规避检测的可能性。我们认为这一场景至关重要,因为它动摇了公众对基准测试可靠性的信任。为更严谨地研究此问题,我们提出对模型提供者及污染检测方法的分类体系。这揭示了现有方法的漏洞,我们利用这些漏洞设计了EAL——一种简单而有效的污染技术,能显著提升基准测试性能,同时完全规避当前的检测方法。