Simulation is an invaluable tool for developing and evaluating controllers for self-driving cars. Current simulation frameworks are driven by highly-specialist domain specific languages, and so a natural language interface would greatly enhance usability. But there is often a gap, consisting of tacit assumptions the user is making, between a concise English utterance and the executable code that captures the user's intent. In this paper we describe a system that addresses this issue by supporting an extended multimodal interaction: the user can follow up prior instructions with refinements or revisions, in reaction to the simulations that have been generated from their utterances so far. We use Large Language Models (LLMs) to map the user's English utterances in this interaction into domain-specific code, and so we explore the extent to which LLMs capture the context sensitivity that's necessary for computing the speaker's intended message in discourse.
翻译:仿真是开发和评估自动驾驶汽车控制器的宝贵工具。当前的仿真框架由高度专业化的领域特定语言驱动,因此自然语言接口将极大地提升可用性。然而,用户简洁的英语表述与捕捉其意图的可执行代码之间通常存在差距,这种差距源于用户未明确表达的隐性假设。本文描述了一个系统,通过支持扩展的多模态交互来解决这一问题:用户可以根据已生成的仿真结果,对先前的指令进行细化或修改。我们使用大型语言模型(LLMs)将用户在此类交互中的英语表述映射为领域特定代码,并由此探究LLMs在多大程度上能够捕捉话语中计算说话者意图所需的上下文敏感性。