This work presents a new task of Text Expansion (TE), which aims to insert fine-grained modifiers into proper locations of the plain text to concretize or vivify human writings. Different from existing insertion-based writing assistance tasks, TE requires the model to be more flexible in both locating and generation, and also more cautious in keeping basic semantics. We leverage four complementary approaches to construct a dataset with 12 million automatically generated instances and 2K human-annotated references for both English and Chinese. To facilitate automatic evaluation, we design various metrics from multiple perspectives. In particular, we propose Info-Gain to effectively measure the informativeness of expansions, which is an important quality dimension in TE. On top of a pre-trained text-infilling model, we build both pipelined and joint Locate&Infill models, which demonstrate the superiority over the Text2Text baselines, especially in expansion informativeness. Experiments verify the feasibility of the TE task and point out potential directions for future research toward better automatic text expansion.
翻译:本文提出了文本扩展(Text Expansion, TE)这一新任务,旨在向普通文本的恰当位置插入细粒度修饰成分,以具体化或生动化人类写作。与现有基于插入的写作辅助任务不同,TE要求模型在定位与生成两方面具备更强的灵活性,同时在保留基本语义方面更加谨慎。我们采用四种互补方法,构建了包含1200万自动生成实例和2000个人工标注参考的数据集,覆盖英文和中文。为促进自动评估,我们从多角度设计了多种度量指标,其中特别提出信息增益(Info-Gain)来有效衡量扩展内容的信息量——这是TE任务中的重要质量维度。基于预训练文本填充模型,我们构建了流水线式和联合式定位与填充(Locate&Infill)模型,相比文本到文本(Text2Text)基线方法展现出显著优势,尤其在扩展信息量方面。实验验证了TE任务的可行性,并为未来实现更优自动文本扩展研究指明了潜在方向。