Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks.
翻译:模型预测结果与预期用途之间的不匹配,会在计算机视觉模型部署时造成不利影响。当任务涉及复杂的结构化输出时,这一问题尤为严重,因为设计能解决这种不匹配的流程会更加困难。在自然语言处理领域,通常采用强化学习技术,通过任务奖励来对齐模型。我们采用这一方法,并证明了其在多个计算机视觉任务(如目标检测、全景分割、着色和图像描述)中具有令人惊讶的有效性。我们相信,这一方法有望被广泛用于更好地让模型与多样化的计算机视觉任务对齐。