Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from visual evidence. Unlike prior studies that attribute this text bias to external factors such as data imbalance or instruction tuning, we propose that the bias originates from the model's internal architecture. Specifically, we hypothesize that visual key vectors (Visual Keys) are out-of-distribution (OOD) relative to the text key space learned during language-only pretraining. Consequently, these visual keys receive systematically lower similarity scores during attention computation, leading to their under-utilization in the context representation. To validate this hypothesis, we extract key vectors from LLaVA and Qwen2.5-VL and analyze their distributional structures using qualitative (t-SNE) and quantitative (Jensen-Shannon divergence) methods. The results provide direct evidence that visual and textual keys occupy markedly distinct subspaces within the attention space. The inter-modal divergence is statistically significant, exceeding intra-modal variation by several orders of magnitude. These findings reveal that text bias arises from an intrinsic misalignment within the attention key space rather than solely from external data factors.
翻译:多模态大语言模型在处理视觉-语言数据时表现出明显的文本输入偏好,这限制了其基于视觉证据进行有效推理的能力。与先前研究将这种文本偏差归因于数据不平衡或指令调优等外部因素不同,我们认为该偏差源于模型的内部架构。具体而言,我们假设视觉键向量相对于纯语言预训练阶段习得的文本键空间而言属于分布外样本。因此,这些视觉键在注意力计算中获得的相似度分数系统性偏低,导致其在上下文表示中被低估。为验证该假设,我们从LLaVA和Qwen2.5-VL中提取键向量,并采用定性(t-SNE)与定量(詹森-香农散度)方法分析其分布结构。结果直接证明,视觉键与文本键在注意力空间中占据显著不同的子空间。跨模态散度具有统计显著性,其量级比模态内差异高出数个数量级。这些发现揭示文本偏差源于注意力键空间的内在错位,而非单纯的外部数据因素。