Data used for analytics and machine learning often take the form of tables with categorical entries. We introduce a family of lossless compression algorithms for such data that proceed in four steps: $(i)$ Estimate latent variables associated to rows and columns; $(ii)$ Partition the table in blocks according to the row/column latents; $(iii)$ Apply a sequential (e.g. Lempel-Ziv) coder to each of the blocks; $(iv)$ Append a compressed encoding of the latents. We evaluate it on several benchmark datasets, and study optimal compression in a probabilistic model for that tabular data, whereby latent values are independent and table entries are conditionally independent given the latent values. We prove that the model has a well defined entropy rate and satisfies an asymptotic equipartition property. We also prove that classical compression schemes such as Lempel-Ziv and finite-state encoders do not achieve this rate. On the other hand, the latent estimation strategy outlined above achieves the optimal rate.
翻译:用于分析与机器学习的表格数据常包含分类条目。本文针对此类数据提出一种无损压缩算法族,其执行步骤如下:$(i)$ 估计与行、列相关的潜变量;$(ii)$ 根据行/列潜变量将表格划分为数据块;$(iii)$ 对每个数据块应用顺序编码器(如Lempel-Ziv算法);$(iv)$ 附加潜变量的压缩编码。我们在多个基准数据集上评估该算法,并在表格数据的概率模型框架下研究了最优压缩——该模型中潜变量相互独立,且表条目在给定潜变量条件下条件独立。我们证明了该模型具有明确的熵率且满足渐近等分性质,同时证明了Lempel-Ziv等经典压缩方案及有限状态编码器无法达到该熵率。另一方面,上述潜变量估计策略可实现最优压缩率。