We present a sequence-to-sequence vision-language model whose parameters are jointly trained on all tasks (all for one) and fully shared among multiple tasks (one for all), resulting in a single model which we named Musketeer. The integration of knowledge across heterogeneous tasks is enabled by a novel feature called Task Explanation Prompt (TEP). TEP reduces interference among tasks, allowing the model to focus on their shared structure. With a single model, Musketeer achieves results comparable to or better than strong baselines trained on single tasks, almost uniformly across multiple tasks.
翻译:我们提出了一种序列到序列的视觉语言模型,其参数在所有任务上联合训练(我为人人),并在多个任务间完全共享(人人为我),最终得到一个单一模型,我们将其命名为“Musketeer”。跨异构任务的知识整合得益于一种名为“任务解释提示”(TEP)的新特性。TEP减少了任务间的干扰,使模型能够聚焦于它们的共享结构。凭借单一模型,Musketeer在多个任务上几乎一致地取得了与强基线模型(在单一任务上训练)相当或更优的结果。