As an interpretable and universal neuro-symbolic paradigm based on Large Language Models, visual programming (VisualProg) can execute compositional visual tasks without training, but its performance is markedly inferior compared to task-specific supervised learning models. To increase its practicality, the performance of VisualProg on specific tasks needs to be improved. However, the non-differentiability of VisualProg limits the possibility of employing the fine-tuning strategy on specific tasks to achieve further improvements. In our analysis, we discovered that significant performance issues in VisualProg's execution originated from errors made by the sub-modules at corresponding visual sub-task steps. To address this, we propose ``VisualProg Distiller", a method of supplementing and distilling process knowledge to optimize the performance of each VisualProg sub-module on decoupled visual sub-tasks, thus enhancing the overall task performance. Specifically, we choose an end-to-end model that is well-performed on the given task as the teacher and further distill the knowledge of the teacher into the invoked visual sub-modules step-by-step based on the execution flow of the VisualProg-generated programs. In this way, our method is capable of facilitating the fine-tuning of the non-differentiable VisualProg frameworks effectively. Extensive and comprehensive experimental evaluations demonstrate that our method can achieve a substantial performance improvement of VisualProg, and outperforms all the compared state-of-the-art methods by large margins. Furthermore, to provide valuable process supervision for the GQA task, we construct a large-scale dataset by utilizing the distillation process of our method.
翻译:作为基于大语言模型的可解释且通用的神经符号范式,视觉编程(VisualProg)无需训练即可执行组合式视觉任务,但其性能显著弱于特定任务的监督学习模型。为提升其实用性,需改进VisualProg在特定任务上的表现。然而,VisualProg的不可微性限制了在特定任务上采用微调策略以进一步改进的可能性。分析发现,VisualProg执行中的显著性能问题源于子模块在对应视觉子任务步骤中的错误。为此,我们提出"VisualProg Distiller"方法,通过补充和蒸馏过程知识来优化各VisualProg子模块在解耦后的视觉子任务上的性能,从而提升整体任务表现。具体而言,我们选择在给定任务上表现优异的端到端模型作为教师,并基于VisualProg生成程序的执行流程,逐步将教师的知识蒸馏到所调用的视觉子模块中。通过这种方式,我们的方法能有效促进不可微VisualProg框架的微调。广泛且全面的实验评估表明,该方法能使VisualProg获得显著的性能提升,并以大幅度优势超越所有对比的先进方法。此外,为GQA任务提供有价值的过程监督,我们利用所提出方法的蒸馏过程构建了一个大规模数据集。