Formatting is an important property in tables for visualization, presentation, and analysis. Spreadsheet software allows users to automatically format their tables by writing data-dependent conditional formatting (CF) rules. Writing such rules is often challenging for users as it requires them to understand and implement the underlying logic. We present FormaT5, a transformer-based model that can generate a CF rule given the target table and a natural language description of the desired formatting logic. We find that user descriptions for these tasks are often under-specified or ambiguous, making it harder for code generation systems to accurately learn the desired rule in a single step. To tackle this problem of under-specification and minimise argument errors, FormaT5 learns to predict placeholders though an abstention objective. These placeholders can then be filled by a second model or, when examples of rows that should be formatted are available, by a programming-by-example system. To evaluate FormaT5 on diverse and real scenarios, we create an extensive benchmark of 1053 CF tasks, containing real-world descriptions collected from four different sources. We release our benchmarks to encourage research in this area. Abstention and filling allow FormaT5 to outperform 8 different neural approaches on our benchmarks, both with and without examples. Our results illustrate the value of building domain-specific learning systems.
翻译:格式化是表格在可视化、展示和分析中的重要属性。电子表格软件允许用户通过编写依赖数据的条件格式化规则来自动格式化表格。由于用户需理解并实现底层逻辑,编写此类规则往往具有挑战性。我们提出FormaT5,一种基于Transformer的模型,能够在给定目标表格和所需格式化逻辑的自然语言描述后生成条件格式化规则。我们发现用户对此类任务的描述常存在规范不足或歧义,导致代码生成系统难以单步准确习得所需规则。为应对这种规范不足问题并最小化参数错误,FormaT5通过弃权目标学习预测占位符。这些占位符可由第二模型填充,或当存在应格式化行的示例时由编程示例系统填充。为在多样化真实场景中评估FormaT5,我们构建了包含1053个条件格式化任务的广泛基准测试集,其中包含从四个不同来源收集的真实世界描述。我们将发布基准测试集以推动该领域研究。弃权与填充机制使FormaT5在基准测试中(无论是否有示例)均优于8种不同神经方法。我们的结果揭示了构建领域专用学习系统的价值。