Internal Language Model (LM)-based methods use permutation language modeling (PLM) to solve the error correction caused by conditional independence in external LM-based methods. However, random permutations of human interference cause fit oscillations in the model training, and Iterative Refinement (IR) operation to improve multimodal information decoupling also introduces additional overhead. To address these issues, this paper proposes the Hierarchical Attention autoregressive Model with Adaptive Permutation (HAAP) to enhance the location-context-image interaction capability, improving autoregressive generalization with internal LM. First, we propose Implicit Permutation Neurons (IPN) to generate adaptive attention masks to dynamically exploit token dependencies. The adaptive masks increase the diversity of training data and prevent model dependency on a specific order. It reduces the training overhead of PLM while avoiding training fit oscillations. Second, we develop Cross-modal Hierarchical Attention mechanism (CHA) to couple context and image features. This processing establishes rich positional semantic dependencies between context and image while avoiding IR. Extensive experimental results show the proposed HAAP achieves state-of-the-art (SOTA) performance in terms of accuracy, complexity, and latency on several datasets.
翻译:基于内部语言模型(LM)的方法采用排列语言建模(PLM)解决外部LM方法中因条件独立性导致的误差校正问题。然而,人为干扰的随机排列会导致模型训练中的拟合震荡,而用于改善多模态信息解耦的迭代精炼(IR)操作也会引入额外开销。为解决这些问题,本文提出具有自适应排列的分层注意力自回归模型(HAAP),以增强位置-上下文-图像交互能力,并通过内部LM提升自回归泛化性能。首先,我们提出隐式排列神经元(IPN)生成自适应注意力掩码,动态挖掘token依赖关系。自适应掩码增加了训练数据的多样性,防止模型对特定顺序产生依赖,在降低PLM训练开销的同时避免训练拟合震荡。其次,我们开发跨模态分层注意力机制(CHA)耦合上下文与图像特征。该处理过程在避免IR操作的同时,建立了上下文与图像之间丰富的位置语义依赖关系。大量实验结果表明,所提出的HAAP在多个数据集上的准确率、复杂度和延迟均达到最优(SOTA)性能。