In this work, we introduce a novel deep learning architecture, Variable Length Embeddings (VLEs), an autoregressive model that can produce a latent representation composed of an arbitrary number of tokens. As a proof of concept, we demonstrate the capabilities of VLEs on tasks that involve reconstruction and image decomposition. We evaluate our experiments on a mix of the iNaturalist and ImageNet datasets and find that VLEs achieve comparable reconstruction results to a state of the art VAE, using less than a tenth of the parameters.
翻译:本文提出了一种新颖的深度学习架构——可变长度嵌入(Variable Length Embeddings, VLEs),这是一种自回归模型,能够生成由任意数量令牌构成的潜在表示。作为概念验证,我们在涉及重建和图像分解的任务中展示了VLEs的能力。我们在iNaturalist和ImageNet混合数据集上进行了实验评估,发现VLEs在仅使用不到最优变分自编码器(VAE)十分之一参数的情况下,达到了与其相当的重建效果。