With the increasing ubiquity of cameras and smart sensors, humanity is generating data at an exponential rate. Access to this trove of information, often covering yet-underrepresented use-cases (e.g., AI in medical settings) could fuel a new generation of deep-learning tools. However, eager data scientists should first provide satisfying guarantees w.r.t. the privacy of individuals present in these untapped datasets. This is especially important for images or videos depicting faces, as their biometric information is the target of most identification methods. While a variety of solutions have been proposed to de-identify such images, they often corrupt other non-identifying facial attributes that would be relevant for downstream tasks. In this paper, we propose Disguise, a novel algorithm to seamlessly de-identify facial images while ensuring the usability of the altered data. Unlike prior arts, we ground our solution in both differential privacy and ensemble-learning research domains. Our method extracts and swaps depicted identities with fake ones, synthesized via variational mechanisms to maximize obfuscation and non-invertibility; while leveraging the supervision from a mixture-of-experts to disentangle and preserve other utility attributes. We extensively evaluate our method on multiple datasets, demonstrating higher de-identification rate and superior consistency than prior art w.r.t. various downstream tasks.
翻译:随着摄像头与智能传感器的日益普及,人类正以指数级速度生成数据。这些数据宝库(尤其是医学影像等尚未充分开发的场景)的获取,有望推动新一代深度学习工具的发展。然而,热切的数据科学家们首先需为这些未开发数据集中涉及的个人隐私提供可靠保障——这对包含人脸信息的图像或视频而言尤为重要,因为生物特征信息是大多数身份识别方法的主要目标。尽管已有多种去识别方案被提出,但它们往往会破坏与下游任务相关的其他非身份性面部属性。本文提出Disguise算法,这是一种能在保证数据实用性的同时实现人脸图像无缝去识别的新型方案。与现有技术不同,我们将解决方案建立在差分隐私与集成学习两大研究领域的基础上。该方法通过提取并替换身份特征为合成假身份(采用变分机制生成以最大化混淆性与不可逆性),同时利用混合专家模型的监督机制分离并保留其他实用属性。我们在多个数据集上进行了全面评估,结果表明该方法在去识别率方面优于现有技术,且针对各类下游任务具有更优越的一致性。