We propose TC-AE, a ViT-based architecture for deep compression autoencoders. Existing methods commonly increase the channel number of latent representations to maintain reconstruction quality under high compression ratios. However, this strategy often leads to latent representation collapse, which degrades generative performance. Instead of relying on increasingly complex architectures or multi-stage training schemes, TC-AE addresses this challenge from the perspective of the token space, the key bridge between pixels and image latents, through two complementary innovations: Firstly, we study token number scaling by adjusting the patch size in ViT under a fixed latent budget, and identify aggressive token-to-latent compression as the key factor that limits effective scaling. To address this issue, we decompose token-to-latent compression into two stages, reducing structural information loss and enabling effective token number scaling for generation. Secondly, to further mitigate latent representation collapse, we enhance the semantic structure of image tokens via joint self-supervised training, leading to more generative-friendly latents. With these designs, TC-AE achieves substantially improved reconstruction and generative performance under deep compression. We hope our research will advance ViT-based tokenizer for visual generation.
翻译:我们提出TC-AE,一种基于ViT的深度压缩自编码器架构。现有方法通常通过增加潜在表示的通道数来在高压缩比下维持重建质量,然而这一策略常导致潜在表示坍缩,从而降低生成性能。TC-AE不依赖于日益复杂的架构或多阶段训练方案,而是从令牌空间(像素与图像潜在表示的关键桥梁)角度通过两项互补创新应对这一挑战:首先,我们研究在固定潜在预算下通过调整ViT中的块大小实现令牌数量缩放,并发现激进的令牌到潜在表示压缩是限制有效缩放的关键因素。为解决该问题,我们将令牌到潜在表示压缩分解为两个阶段,减少结构信息损失并实现有效的令牌数量缩放以支持生成。其次,为进一步缓解潜在表示坍缩,我们通过联合自监督训练增强图像令牌的语义结构,从而获得更利于生成的潜在表示。基于这些设计,TC-AE在深度压缩条件下实现了显著提升的重建与生成性能。我们期望此项研究能推进基于ViT的标记器在视觉生成领域的应用。