In this paper, we introduce a metadata-enriched generation framework (PhishFuzzer) that seeds real emails into Large Language Models (LLMs) to produce 23,100 diverse, structurally consistent email variants across controlled entity and length dimensions. Unlike prior corpora, our dataset features strict three-class labels (Phishing, Spam, Valid), provides full URL and attachment metadata, and annotates each email with attacker intent. Using this dataset, we benchmark two state-of-the-art LLMs (Qwen-2.5-72B and Gemini-3.1-Pro) under both Basic (body, subject) and Full (+URL, sender, attachment) settings. By applying formal confidence metrics (Task Success Rate and Confidence Index), we analyze model reliability, robustness against linguistic fuzzing, and the impact of structural metadata on detection accuracy. Our fully open-source framework and dataset provide a rigorous foundation for evaluating next-generation email security systems. To support open science, we make the PhishFuzzer Dataset, the generation scripts and prompts available on GitHub: https://github.com/DataPhish/PhishFuzzer
翻译:本文提出一种元数据增强的生成框架(PhishFuzzer),通过将真实邮件注入大语言模型(LLMs),在可控实体与长度维度上生成23,100个多样化且结构一致的邮件变体。与现有语料库不同,本数据集采用严格的三分类标签(钓鱼/垃圾/正常),提供完整的URL与附件元数据,并为每封邮件标注攻击者意图。基于该数据集,我们在基础设置(正文、主题)与完整设置(+URL、发件人、附件)下对两种先进大语言模型(Qwen-2.5-72B和Gemini-3.1-Pro)进行基准测试。通过引入形式化置信度指标(任务成功率与置信度指数),我们分析了模型可靠性、对语言模糊化的鲁棒性,以及结构元数据对检测准确率的影响。本框架与数据集完全开源,为评估下一代邮件安全系统提供了严谨基础。为支持开放科学,我们在GitHub上公开了PhishFuzzer数据集、生成脚本及提示信息:https://github.com/DataPhish/PhishFuzzer