Generative tabular augmentation is appealing in data-scarce domains, yet the prevailing focus on distributional fidelity does not reliably translate into better downstream models. We formalize a fidelity-utility gap: common generative objectives prioritize distributional plausibility, whereas augmentation succeeds only when injected samples reduce the current learner's held-out evaluation loss. This gap motivates learning not just how to generate, but what to generate and when to inject as training evolves. We propose TAP (Tabular Augmentation Policy), which couples diffusion inpainting with a lightweight, learner-conditioned policy to steer generation toward high-utility regions and controls safe injection via explicit gating and conservative windowed commitment. Under severe data scarcity, TAP consistently outperforms strong generative baselines on seven real-world datasets, improving classification accuracy by up to 15.6 percentage points and reducing regression RMSE by up to 32%.
翻译:生成式表格数据增强在数据稀缺领域具有吸引力,然而当前对分布保真度的侧重并不能可靠地转化为更优的下游模型。我们形式化了保真度-效用差距:常见的生成目标优先考虑分布合理性,而数据增强只有在注入样本降低当前学习器保留集评估损失时才成功。这一差距表明需要学习的不仅是"如何生成",还有"生成什么"以及"在训练演化中何时注入"。我们提出TAP(表格数据增强策略),该方法将扩散修复与轻量级学习器条件策略相结合,引导生成向高效用区域发展,并通过显式门控与保守窗口提交机制控制安全注入。在严重数据稀缺条件下,TAP在七个真实世界数据集上持续优于强生成基线,分类准确率提升高达15.6个百分点,回归RMSE降低高达32%。