Large language models are widespread, with their performance on benchmarks frequently guiding user preferences for one model over another. However, the vast amount of data these models are trained on can inadvertently lead to contamination with public benchmarks, thus compromising performance measurements. While recently developed contamination detection methods try to address this issue, they overlook the possibility of deliberate contamination by malicious model providers aiming to evade detection. We argue that this setting is of crucial importance as it casts doubt on the reliability of public benchmarks. To more rigorously study this issue, we propose a categorization of both model providers and contamination detection methods. This reveals vulnerabilities in existing methods that we exploit with EAL, a simple yet effective contamination technique that significantly inflates benchmark performance while completely evading current detection methods.
翻译:大型语言模型已广泛普及,其基准测试性能常常引导用户对不同模型的偏好选择。然而,这些模型训练所依赖的海量数据可能无意中导致与公共基准测试的数据污染,进而损害性能评估的可靠性。尽管近期发展的污染检测方法试图解决这一问题,但它们忽略了恶意模型提供者有意进行污染以规避检测的可能性。我们认为这一场景至关重要,因为它对公共基准测试的可信度提出了质疑。为更严谨地研究该问题,我们提出了模型提供者与污染检测方法的分类体系。这揭示了现有方法中的漏洞,并据此提出EAL技术——一种简单而有效的污染手法,能在完全规避当前检测方法的同时显著提升基准测试性能。