Cross-task generalization is a significant outcome that defines mastery in natural language understanding. Humans show a remarkable aptitude for this, and can solve many different types of tasks, given definitions in the form of textual instructions and a small set of examples. Recent work with pre-trained language models mimics this learning style: users can define and exemplify a task for the model to attempt as a series of natural language prompts or instructions. While prompting approaches have led to higher cross-task generalization compared to traditional supervised learning, analyzing 'bias' in the task instructions given to the model is a difficult problem, and has thus been relatively unexplored. For instance, are we truly modeling a task, or are we modeling a user's instructions? To help investigate this, we develop LINGO, a novel visual analytics interface that supports an effective, task-driven workflow to (1) help identify bias in natural language task instructions, (2) alter (or create) task instructions to reduce bias, and (3) evaluate pre-trained model performance on debiased task instructions. To robustly evaluate LINGO, we conduct a user study with both novice and expert instruction creators, over a dataset of 1,616 linguistic tasks and their natural language instructions, spanning 55 different languages. For both user groups, LINGO promotes the creation of more difficult tasks for pre-trained models, that contain higher linguistic diversity and lower instruction bias. We additionally discuss how the insights learned in developing and evaluating LINGO can aid in the design of future dashboards that aim to minimize the effort involved in prompt creation across multiple domains.
翻译:跨任务泛化是定义自然语言理解掌握程度的重要成果。人类在此方面表现出卓越的能力,能够根据文本指令形式的定义和少量示例解决多种不同类型的任务。近期基于预训练语言模型的研究模仿了这种学习方式:用户可通过一系列自然语言提示或指令来定义和举例说明模型需尝试的任务。尽管提示方法相比传统监督学习实现了更强的跨任务泛化能力,但分析输入模型的任务指令中的"偏差"仍是一个难题,相关研究相对较少。例如,我们究竟是在建模真实任务,还是在建模用户的指令?为探究此问题,我们开发了LINGO——一种新颖的可视分析界面,支持高效的任务驱动工作流:(1)帮助识别自然语言任务指令中的偏差,(2)修改(或创建)任务指令以降低偏差,(3)评估预训练模型在去偏任务指令上的性能。为全面评估LINGO,我们面向包含55种语言的1616个语言任务及其自然语言指令的数据集,对新手与专家指令创建者开展了用户研究。结果表明,两组用户均能通过LINGO为预训练模型创建更具难度的任务,这些任务包含更高的语言多样性和更低的指令偏差。此外,我们探讨了LINGO开发与评估中获得的见解如何指导未来仪表板设计,以最小化跨领域提示创建所需的工作量。