Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diffusion models remain brittle. Existing benchmarks evaluate only one axis at a time and none jointly captures multi-identity composition with human-object interaction, background grounding, and spatial plausibility. We introduce CogCanvas, a benchmark of 1,952 curated reference images spanning 100 celebrity identities, 115 distinctive objects and fashion items, and 29 real-world background scenes including landmarks, from which we construct 1,361 compositional prompts covering 2-5 person group sizes. The curation pipeline combines DINOv2-based deduplication, two-stage aesthetic filtering, and automated derivation of structured interaction and position graphs that serve as ground-truth supervision. CogCanvas supports three tasks, reference-based multi-human-object generation (primary), text-to-image compositional generation, and reference retrieval, under a unified six-axis evaluation protocol. We introduce two metrics tailored to the multi-reference setting: BG-Sim, which scores background fidelity on SAM 3-masked regions via DINOv3 feature similarity, and Attr-VQA, which uses a multimodal LLM to verify per-subject attribute binding and inter-person interactions against the structured graphs. Benchmarking five SOTA methods reveals that every model degrades substantially as group size grows from 2 to 5, with near-complete failure on object/fashion binding beyond three subjects.
翻译:多主体参考图像生成需同时保留多个人物身份、绑定每个个体关联的物体与时尚物品,并匹配指定的背景场景,当前扩散模型在此任务中仍显脆弱。现有基准仅评估单一维度,未能同时捕捉含人-物交互的多身份组合、背景接地及空间合理性。本文提出CogCanvas基准,包含1,952张精选参考图像,覆盖100个名人身份、115个独特物体与时尚物品,以及29个含地标的真实背景场景,据此构建1,361条组合提示,涵盖2-5人群体规模。数据筛选流程结合基于DINOv2的去重、两阶段美学过滤,以及结构化交互图与位置图的自动推导,作为真值监督。CogCanvas在统一六轴评估协议下支持三项任务:基于参考的多人物-物体生成(主任务)、文本到图像组合生成及参考检索。我们引入两个针对多参考场景的指标:BG-Sim通过DINOv3特征相似度评估SAM三掩码区域的背景保真度;Attr-VQA利用多模态大语言模型基于结构化图验证各主体属性绑定与人际交互。对五种最新方法的基准测试表明,所有模型在群体规模从2增至5时性能均显著下降,且当主体超过三个时,物体与时尚物品绑定几乎完全失效。