Recent advances in text-to-image synthesis make it possible to visualize machine imaginations for a given context. On the other hand, when generating text, human writers are gifted at creative visualization, which enhances their writings by forming imaginations as blueprints before putting down the stories in words. Inspired by such a cognitive process, we ask the natural question of whether we can endow machines with the same ability to utilize visual information and construct a general picture of the context to guide text generation. In this work, we propose iNLG that uses machine-generated images to guide language models in open-ended text generation. The experiments and analyses demonstrate the effectiveness of iNLG on open-ended text generation tasks, including text completion, story generation, and concept-to-text generation in both few-shot and full-data scenarios. Both automatic metrics and human evaluations verify that the text snippets generated by our iNLG are coherent and informative while displaying minor degeneration.
翻译:近期文本到图像合成的进展使得机器能够针对给定语境进行可视化想象。另一方面,人类写作者在生成文本时,天生擅长创造性可视化——在将故事付诸文字之前,先通过形成想象蓝图来增强写作质量。受此类认知过程的启发,我们提出一个自然问题:能否赋予机器相同的能力,使其利用视觉信息构建语境的整体图景来引导文本生成?本研究提出了iNLG,该方法利用机器生成的图像引导语言模型进行开放式文本生成。实验与分析表明,iNLG在文本补全、故事生成以及概念到文本生成等开放式文本生成任务中均具有有效性,涵盖少样本和全数据场景。自动指标与人工评估均验证了iNLG生成的文本片段具有连贯性和信息性,且仅表现出轻微退化现象。