Visual Question Answering (VQA) using multi-modal data facilitates real-life applications, such as home robots and medical diagnoses. However, one significant challenge is to design a robust learning method for various client tasks. One critical aspect is to ensure privacy, as client data sharing is limited due to confidentiality concerns. This work focuses on addressing the issue of confidentiality constraints in multi-client VQA tasks and limited labeled training data of clients. We propose the Unidirectional Split Learning with Contrastive Loss (UniCon) method to overcome these limitations. The proposed method trains a global model on the entire data distribution of different clients, learning refined cross-modal representations through model sharing. Privacy is ensured by utilizing a split learning architecture in which a complete model is partitioned into two components for independent training. Moreover, recent self-supervised learning techniques were found to be highly compatible with split learning. This combination allows for rapid learning of a classification task without labeled data. Furthermore, UniCon integrates knowledge from various local tasks, improving knowledge sharing efficiency. Comprehensive experiments were conducted on the VQA-v2 dataset using five state-of-the-art VQA models, demonstrating the effectiveness of UniCon. The best-performing model achieved a competitive accuracy of 49.89%. UniCon provides a promising solution to tackle VQA tasks in a distributed data silo setting while preserving client privacy.
翻译:视觉问答(VQA)利用多模态数据可促进家庭机器人和医疗诊断等实际应用。然而,一个关键挑战在于为多样化的客户端任务设计鲁棒的学习方法。核心问题之一是确保隐私性——由于保密限制,客户端数据共享十分有限。本工作聚焦于解决多客户端VQA任务中的保密约束问题以及客户端标注训练数据不足的难题。我们提出单向分割学习与对比损失(UniCon)方法以克服这些限制。所提方法在不同客户端的全数据分布上训练全局模型,通过模型共享学习精细化的跨模态表征。通过采用分割学习架构将完整模型划分为两个独立训练的组件,从而保障隐私性。此外,近期自监督学习技术被发现与分割学习高度兼容,这种结合使得无需标注数据即可快速学习分类任务。更进一步,UniCon整合了来自多种本地任务的知识,提升了知识共享效率。我们基于VQA-v2数据集,采用五种最先进的VQA模型进行了全面实验,验证了UniCon的有效性。性能最优模型取得了49.89%的竞争性准确率。UniCon为在分布式数据孤岛环境下处理VQA任务且保护客户端隐私提供了一种有前景的解决方案。