We are witnessing a novel era of creativity where anyone can create digital content via prompt-based learning (known as prompt engineering). This paper delves into prompt engineering as a novel creative skill for creating AI art with text-to-image generation. In a pilot study, we find that many crowdsourced participants have knowledge about art which could be used for writing effective prompts. In three subsequent studies, we explore whether crowdsourced participants can put this knowledge into practice. We examine if participants can 1) discern prompt quality, 2) write prompts, and 3) refine prompts. We find that participants could evaluate prompt quality and crafted descriptive prompts, but they lacked style-specific vocabulary necessary for effective prompting. This is in line with our hypothesis that prompt engineering is a new type of skill that is non-intuitive and must first be acquired (e.g., through means of practice and learning) before it can be used. Our studies deepen our understanding of prompt engineering and chart future research directions. We offer nine guidelines for conducting research on text-to-image generation and prompt engineering with paid crowds. We conclude by envisioning four potential futures for prompt engineering.
翻译:我们正见证一个创意新时代的诞生——任何人都能通过基于提示的学习(即提示工程)创作数字内容。本文深入探究提示工程作为一项运用文本到图像生成技术创作AI艺术的新兴创意技能。在一项初步研究中,我们发现众多众包参与者具备可用于撰写有效提示的艺术知识。在后续三项研究中,我们考察众包参与者能否将这种知识付诸实践:检验参与者能否1) 辨别提示质量,2) 撰写提示,以及3) 优化提示。研究结果表明,参与者能够评估提示质量并创作描述性提示,但缺乏有效提示所需的风格特异性词汇。这印证了我们的假设:提示工程是一种新型技能,具有非直觉性特征,必须通过实践和学习等手段先掌握才能运用。本研究深化了对提示工程的理解,并指明了未来研究方向。我们提出了九项指导原则,用于指导基于付费众包的文本到图像生成及提示工程研究。最后,我们展望了提示工程的四种潜在发展方向。