This paper highlights the need to bring document classification benchmarking closer to real-world applications, both in the nature of data tested ($X$: multi-channel, multi-paged, multi-industry; $Y$: class distributions and label set variety) and in classification tasks considered ($f$: multi-page document, page stream, and document bundle classification, ...). We identify the lack of public multi-page document classification datasets, formalize different classification tasks arising in application scenarios, and motivate the value of targeting efficient multi-page document representations. An experimental study on proposed multi-page document classification datasets demonstrates that current benchmarks have become irrelevant and need to be updated to evaluate complete documents, as they naturally occur in practice. This reality check also calls for more mature evaluation methodologies, covering calibration evaluation, inference complexity (time-memory), and a range of realistic distribution shifts (e.g., born-digital vs. scanning noise, shifting page order). Our study ends on a hopeful note by recommending concrete avenues for future improvements.}
翻译:本文强调将文档分类基准测试更贴近实际应用的必要性,既涉及测试数据的本质属性($X$:多通道、多页面、多行业;$Y$:类别分布与标签集多样性),也涵盖所考虑的各类分类任务($f$:多页面文档分类、页面流分类、文档包分类等)。我们识别出公共多页面文档分类数据集的匮乏问题,系统形式化应用场景中出现的不同分类任务,并论证了面向高效多页面文档表示的研究价值。基于所提出的多页面文档分类数据集的实验研究表明,当前基准测试已不再适用,需更新以评估实践中自然出现的完整文档。这一现实检验同样呼吁更为成熟的评估方法论,涵盖标定评估、推理复杂度(时间-内存)及一系列现实分布偏移(如原生数字文档与扫描噪声的差异、页面顺序变更)。我们的研究以乐观展望收尾,为未来改进提供了具体方向。