Hardware security competitions such as HackTheSilicon serve as benchmarking platforms for evaluating vulnerability detection methods and for training humans and AI. However, our study reveals that LLMs threaten their validity. Instead of genuine security reasoning, detectors exploit a diff-style syntactic comparison, achieving an 83% detection rate, undermining fair evaluation. To mitigate this, we propose the first LLM-oriented, semantics-preserving obfuscation framework for these benchmarks. Unlike IP-protection approaches, it applies human-readable transformations and controlled diff-noise while preserving functionality. On HackTheSilicon, the framework reduces LLM-based detection accuracy by 50% with only 10% obfuscation and by 78.6% under complete obfuscation, restoring benchmark reliability.
翻译:硬件安全竞赛(如HackTheSilicon)作为评估漏洞检测方法及训练人类与人工智能的基准平台,发挥着重要作用。然而,本研究表明,大语言模型(LLM)正威胁着这些平台的有效性:检测器并非基于真正的安全推理,而是利用类似diff的句法比较方式,实现了高达83%的检测率,严重破坏了公平评估机制。为此,我们首次提出面向LLM、语义保持的混淆框架。与知识产权保护方案不同,该框架在保持功能完整性的前提下,采用人类可读的变换方式并引入受控的diff噪声。在HackTheSilicon平台上,该框架可使基于LLM的检测准确率降低50%(仅需10%混淆度),并在完全混淆条件下降低78.6%,从而恢复基准测试的可靠性。