Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, text, timestamps, locations, and repeated events. Existing visual, video, document, and lifelog benchmarks test sub-problems, but not album-scale profile reconstruction with social identity binding and evidence citation. Benchmarking this task is difficult because the ground truth needed for evaluation--owner profiles, social graphs, face-name maps, and evidence provenance--is private state that real albums cannot safely release. We introduce PAL-Bench, a controlled benchmark for evidence-grounded reconstruction under a public-record contract. Its Evidence Compiler builds latent private worlds, programs target-level evidence paths, renders album pixels, re-measures them through perception pipelines, and exports audited public/private views. Agents receive only perception-derived public records; targets, identifier maps, and evidence paths remain hidden. PAL-Bench contains 50 synthetic users, 36,659 public photo records, and 2,799 targets over owner facts, identities, and relations. A privacy-preserving audit with 10 participants confirms that PAL-Bench evidence structures match real private albums, though equivalent releases remain privacy-prohibitive. Across seven systems and two compute-matched diagnostics, a seven-metric protocol reveals a gap between plausible profile summarization and faithful social reconstruction: systems recover some owner facts but struggle with recurring identities and evidence citation. PAL-TRACE, a reference framework that freezes identity bindings before owner-fact mining, performs best but leaves hard identity resolution far from solved. PAL-Bench provides a testbed for perceptual entity resolution, multimodal data integration, temporal evidence aggregation, and provenance-aware structured prediction.
翻译:纵向个人相册是弱模式多模态数据库:包含噪声感知记录,其关键事实需要跨人脸、文本、时间戳、位置和重复事件进行连接。现有视觉、视频、文档和生活日志基准测试仅解决子问题,而非涵盖社交身份绑定与证据引用的相册级画像重建。此类基准测试的难点在于:用于评估的真值数据(所有者画像、社交图谱、人脸-姓名映射及证据溯源)属于真实相册无法安全公开的隐私状态。我们提出PAL-Bench——一项基于公共记录契约的可控基准测试,用于证据驱动型重建。其证据编译器构建潜在隐私世界、编程目标级证据路径、渲染相册像素、通过感知管线重新测量像素,并导出经审计的公共/隐私视图。智能体仅能获取感知派生的公共记录,而目标、标识映射及证据路径保持隐藏。PAL-Bench包含50名合成用户、36,659张公共照片记录及2,799个涉及所有者事实、身份与关系的目标。经10名参与者参与的隐私保护审计证实,PAL-Bench的证据结构与真实私人相册一致(尽管同等规模的数据发布仍受隐私限制)。在七个系统与两项计算匹配诊断中,七指标评估协议揭示了“合理的画像摘要”与“真实社交重建”之间的鸿沟:系统能恢复部分所有者事实,但在重复身份识别与证据引用方面表现欠佳。参考框架PAL-TRACE(在挖掘所有者事实前冻结身份绑定)表现最佳,但硬身份解析问题远未解决。PAL-Bench为感知实体解析、多模态数据集成、时序证据聚合及溯源感知结构化预测提供了测试平台。