A model can learn that the piano piece Für Elise is calm and reflective by listening to the audio or by reading a text description, but does it matter which route that knowledge took when it is later at risk of being forgotten? Forgetting research in multimodal models measures what knowledge is lost under adaptation, yet has not asked whether acquisition route affects how easily that knowledge is forgotten. We call this untested premise the Pathway-Invariant Assumption. Music understanding enables a clean test because a music clip and a canonical text description can be aligned to the same perceptual content, allowing the same knowledge unit to enter a model through listening or reading while the target remains fixed. Across multiple architecturally distinct audio-language models, we observe a consistent asymmetry: text-pathway knowledge is forgotten more than matched audio-pathway knowledge under identical adaptation pressure. To attribute this effect to route rather than confounds, we introduce the Paired Pathway Controlled Protocol (PPCP), a three-phase design that establishes matched pathway baselines, activates both pathways under symmetric supervision on the same knowledge pool, and applies identical forgetting pressure to both pathways. The gap is stable across models and gain-controlled analyses, persists when contradictory overwrite is replaced by correct-label cross-domain learning, remains under single-modality pressure, and is not removed by lightweight replay. Two independent routing-depth controls confirm that the effect is not explained by architectural depth, pointing to input representation as the dominant factor. Under PPCP, our results demonstrate that forgetting is highly route-dependent, establishing acquisition route as a new analytical dimension for forgetting research and multimodal system design.
翻译:模型可以通过聆听音频或阅读文字描述来学习钢琴曲《致爱丽丝》的宁静与沉思,但当这些知识日后面临遗忘风险时,获取路径是否会影响遗忘的难易程度?多模态模型中的遗忘研究关注的是在适配过程中哪些知识会丢失,但尚未探究获取路径是否会影响知识被遗忘的难易程度。我们将这一未被检验的前提称为"路径不变假设"。音乐理解为此提供了理想的检验场景,因为一段音乐片段与一段规范的文字描述可以对应相同的感知内容,使得同一知识单元能够通过聆听或阅读进入模型,同时目标保持不变。在多个架构各异的音频-语言模型中,我们观察到一致的非对称性:在相同的适配压力下,通过文本路径获取的知识比通过匹配的音频路径获取的知识更容易被遗忘。为将这一效应归因于路径而非混淆因素,我们引入了"配对路径受控协议"(PPCP),这是一种三阶段设计:建立匹配的路径基线,在对称监督下通过同一知识池激活两条路径,并对两条路径施加相同的遗忘压力。该差距在不同模型和增益控制分析中保持稳定,在将矛盾覆盖替换为正确标签的跨领域学习时依然存在,在单模态压力下持续出现,且无法通过轻量级重放消除。两项独立的路径深度控制实验证实,该效应不能由架构深度解释,输入表征才是主导因素。在PPCP框架下,我们的结果表明遗忘高度依赖于路径,从而确立了获取路径作为遗忘研究与多模态系统设计的新分析维度。