Although machine unlearning is essential for removing private, harmful, or copyrighted content from LLMs, current benchmarks often fail to faithfully represent the true ``forgetting scope'' learned by the model. We formalize two distinct unlearning granularities, domain-level and instance-level, and propose \BiForget, an automated framework for synthesizing high-quality forget sets. Unlike prior work relying on \emph{external} generators, \BiForget exploits the target model per se to elicit data that matches its internal knowledge distribution through seed-guided and adversarial prompting. Our experiments across diverse benchmarks show that it achieves a superior balance of relevance, diversity, and efficiency. Quantitatively, in the Harry Potter domain, it improves relevance by ${\sim}20$ and diversity by ${\sim}$0.05 while \emph{halving} the total data size compared to SOTAs. Ultimately, it facilitates more robust forgetting and better utility preservation, providing a more rigorous foundation for evaluating LLM unlearning.
翻译:尽管机器遗忘对于从大语言模型中移除隐私、有害或受版权保护的内容至关重要,但现有基准测试往往无法真实反映模型实际习得的"遗忘范围"。我们形式化定义了两种不同的遗忘粒度——领域级与实例级,并提出双遗忘框架(\BiForget),一种自动化合成高质量遗忘集合的方法。与依赖外部生成器的现有工作不同,\BiForget 通过种子引导与对抗性提示,直接利用目标模型自身来调用与其内部知识分布相匹配的数据。我们在多个基准测试上的实验表明,该方法在相关性、多样性与效率之间实现了卓越平衡。定量来看,在哈利波特领域,相比当前最优方法,其在将总数据量减半的同时,相关性提升了约20,多样性提升了约0.05。最终,该方法促进了更鲁棒的遗忘与更好的效用保持,为评估大语言模型遗忘能力提供了更严谨的基础。