Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redundancy or information loss. Drawing inspiration from Shannon's information theory, this paper introduces InfoTok, a principled framework for adaptive video tokenization. We rigorously prove that existing data-agnostic training methods are suboptimal in representation length, and present a novel evidence lower bound (ELBO)-based algorithm that approaches theoretical optimality. Leveraging this framework, we develop a transformer-based adaptive compressor that enables adaptive tokenization. Empirical results demonstrate state-of-the-art compression performance, saving 20% tokens without influence on performance, and achieving 2.3x compression rates while still outperforming prior heuristic adaptive approaches. By allocating tokens according to informational richness, InfoTok enables a more compressed yet accurate tokenization for video representation, offering valuable insights for future research.
翻译:精确且高效的离散视频分词对长视频序列处理至关重要。然而,视频固有的复杂性和可变信息密度对当前分词器造成了显著瓶颈——这些分词器以固定速率机械地压缩所有内容,导致冗余或信息丢失。受香农信息论启发,本文引入InfoTok,一种用于自适应视频分词的原则性框架。我们严格论证了现有与数据无关的训练方法在表示长度上并非最优,并提出一种基于证据下界(ELBO)的新算法,逼近理论最优性。借助此框架,我们开发了基于Transformer的自适应压缩器实现自适应分词。实验结果表明,该方法达到了最先进的压缩性能:在性能不受影响的情况下节省20%的token,且实现2.3倍压缩率的同时仍优于先前的启发式自适应方法。通过根据信息丰富度分配token,InfoTok实现了更紧凑且精确的视频表示分词,为未来研究提供了宝贵见解。