In this paper, we introduce LLaVA-$\phi$ (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate multi-modal dialogues. LLaVA-Phi marks a notable advancement in the realm of compact multi-modal models. It demonstrates that even smaller language models, with as few as 2.7B parameters, can effectively engage in intricate dialogues that integrate both textual and visual elements, provided they are trained with high-quality corpora. Our model delivers commendable performance on publicly available benchmarks that encompass visual comprehension, reasoning, and knowledge-based perception. Beyond its remarkable performance in multi-modal dialogue tasks, our model opens new avenues for applications in time-sensitive environments and systems that require real-time interaction, such as embodied agents. It highlights the potential of smaller language models to achieve sophisticated levels of understanding and interaction, while maintaining greater resource efficiency.The project is available at {https://github.com/zhuyiche/llava-phi}.
翻译:本文提出LLaVA-$\phi$(LLaVA-Phi)——一种高效的多模态助手,利用近期先进的小语言模型Phi-2促进多模态对话。LLaVA-Phi标志着紧凑型多模态模型领域的重要进展,证明即使仅有27亿参数的小语言模型,只要经过高质量语料训练,也能有效参与融合文本与视觉元素的复杂对话。我们的模型在包含视觉理解、推理与知识感知的公开基准测试中展现出优越性能。除在多模态对话任务中的突出表现外,该模型为时间敏感环境与需要实时交互的系统(如具身智能体)开辟了新应用途径。研究凸显了小语言模型在保持更高资源效率的同时,实现复杂理解与交互能力的潜力。项目地址:https://github.com/zhuyiche/llava-phi。