Test stimuli generation has been a crucial but labor-intensive task in hardware design verification. In this paper, we revolutionize this process by harnessing the power of large language models (LLMs) and present a novel benchmarking framework, LLM4DV. This framework introduces a prompt template for interactively eliciting test stimuli from the LLM, along with four innovative prompting improvements to support the pipeline execution and further enhance its performance. We compare LLM4DV to traditional constrained-random testing (CRT), using three self-designed design-under-test (DUT) modules. Experiments demonstrate that LLM4DV excels in efficiently handling straightforward DUT scenarios, leveraging its ability to employ basic mathematical reasoning and pre-trained knowledge. While it exhibits reduced efficiency in complex task settings, it still outperforms CRT in relative terms. The proposed framework and the DUT modules used in our experiments will be open-sourced upon publication.
翻译:测试激励生成一直是硬件设计验证中至关重要但劳动密集型的任务。本文通过利用大型语言模型(LLM)的力量彻底革新了这一过程,并提出了一种新颖的基准测试框架LLM4DV。该框架引入了一种提示模板,用于交互式地从LLM中获取测试激励,同时提供了四项创新性的提示改进措施,以支持流水线执行并进一步提升其性能。我们将LLM4DV与传统的约束随机测试(CRT)进行了比较,使用了三个自行设计的待测设计(DUT)模块。实验表明,LLM4DV凭借其运用基础数学推理和预训练知识的能力,在高效处理简单DUT场景方面表现出色。虽然在复杂任务设置中其效率有所降低,但在相对意义上仍优于CRT。本文提出的框架及实验中使用的DUT模块将在发表后开源。