This paper studies ensembling in the era of Large Vision-Language Models (LVLMs). Ensembling is a classical method to combine different models to get increased performance. In the recent work on Encyclopedic-VQA the authors examine a wide variety of models to solve their task: from vanilla LVLMs, to models including the caption as extra context, to models augmented with Lens-based retrieval of Wikipedia pages. Intuitively these models are highly complementary, which should make them ideal for ensembling. Indeed, an oracle experiment shows potential gains from 48.8% accuracy (the best single model) all the way up to 67% (best possible ensemble). So it is a trivial exercise to create an ensemble with substantial real gains. Or is it?
翻译:本文研究了大型视觉-语言模型(LVLMs)时代中的集成方法。集成是一种经典方法,通过组合不同模型来提升性能。在近期关于百科全书式视觉问答(Encyclopedic-VQA)的研究中,作者考察了多种模型以解决其任务:从原始LVLMs到将图像描述作为额外上下文的模型,再到基于Lens检索维基百科页面增强的模型。直觉上,这些模型高度互补,使其成为集成方法的理想候选。事实上,一项或然实验显示,潜在收益可从48.8%的准确率(最佳单模型)提升至67%(最佳可能集成)。因此,构建一个能带来实质性收益的集成似乎是轻而易举的。但事实果真如此吗?