This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around a one-level Haar DWT/IDWT frontend, a shared coefficient-token layout, optional structural metadata, lightweight modality value adapters, and a shared token-wise encoder-decoder trunk. On Speech Commands, EuroSAT RGB, and DAVIS 2017 data, a dense shared model reaches 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR. A matched-rate sweep under continuous latent scalar budgets indicates that the visual gains are not explained solely by latent capacity, while also showing that additive metadata embeddings are not a universal source of improvement. Finally, fixed-rate energy selection provides a strong non-parametric baseline: energy_global improves average PSNR over uniform selection by 16.73 dB for audio, 16.90 dB for images, and 15.86 dB for video under compressed keep ratios. Masked sparse training reaches 34.45 dB video PSNR with 50% of dense tokens. The results support a unified wavelet token schema and sparse token interface, while stopping short of establishing a universal discrete vocabulary.
翻译:本文研究音频、图像和视频是否可以共享一种通用的小波标记模式,而非依赖各自模态特有的潜在网格。我们介绍了一种初步的连续令牌模型,其核心组件包括:单级Haar DWT/IDWT前端、共享系数令牌布局、可选的结构元数据、轻量级模态值适配器,以及共享的令牌级编解码主干。在Speech Commands、EuroSAT RGB和DAVIS 2017数据集上,密集共享模型分别达到39.92 dB(音频)、29.37 dB(图像)和23.93 dB(视频)的PSNR。在连续潜在标量预算下的匹配速率扫描表明,视觉增益并非仅由潜在容量解释,同时显示加性元数据嵌入并非普遍的改进来源。最后,固定速率能量选择提供了一个强大的非参数基线:在压缩保留比下,energy_global相比均匀选择将音频、图像和视频的平均PSNR分别提升了16.73 dB、16.90 dB和15.86 dB。掩码稀疏训练在50%密集令牌比例下达34.45 dB视频PSNR。这些结果支持了统一的小波令牌模式和稀疏令牌接口,但尚不足以建立通用的离散词汇表。