See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action Designers

Automated vehicles lack natural communication channels with other road users, making external Human-Machine Interfaces (eHMIs) essential for conveying intent and maintaining trust in shared environments. However, most eHMI studies rely on developer-crafted message-action pairs, which are difficult to adapt to diverse and dynamic traffic contexts. A promising alternative is to use Large Language Models (LLMs) as action designers that generate context-conditioned eHMI actions, yet such designers lack perceptual verification and typically depend on fixed prompts or costly human-annotated feedback for improvement. We present See2Refine, a human-free, closed-loop framework that uses vision-language model (VLM) perceptual evaluation as automated visual feedback to improve an LLM-based eHMI action designer. Given a driving context and a candidate eHMI action, the VLM evaluates the perceived appropriateness of the action, and this feedback is used to iteratively revise the designer's outputs, enabling systematic refinement without human supervision. We evaluate our framework across three eHMI modalities (lightbar, eyes, and arm) and multiple LLM model sizes. Across settings, our framework consistently outperforms prompt-only LLM designers and manually specified baselines in both VLM-based metrics and human-subject evaluations. Results further indicate that the improvements generalize across modalities and that VLM evaluations are well aligned with human preferences, supporting the robustness and effectiveness of See2Refine for scalable action design.

翻译：自动驾驶车辆缺乏与其他道路使用者的自然沟通渠道，这使得外部人机界面（eHMI）对于在共享环境中传达意图和维持信任至关重要。然而，大多数eHMI研究依赖于开发者预先定义的消息-行为配对，难以适应多样且动态的交通场景。一种有前景的替代方案是使用大型语言模型（LLM）作为行为设计器来生成基于情境的eHMI行为，但此类设计器缺乏感知验证，通常依赖固定提示或成本高昂的人工标注反馈进行改进。我们提出See2Refine——一个无需人工介入的闭环框架，该框架利用视觉-语言模型（VLM）的感知评估作为自动化视觉反馈，以改进基于LLM的eHMI行为设计器。给定驾驶情境和候选eHMI行为，VLM评估该行为在感知层面的适宜性，并利用该反馈迭代修正设计器的输出，从而在无需人工监督的情况下实现系统性优化。我们在三种eHMI模态（灯带、眼睛和手臂）及多种LLM模型规模下评估本框架。在所有实验设置中，我们的框架在基于VLM的指标和人类受试者评估中均持续优于仅使用提示的LLM设计器及人工设定的基线方法。结果进一步表明，性能提升在不同模态间具有泛化性，且VLM评估与人类偏好高度一致，这证明了See2Refine框架在可扩展行为设计方面的鲁棒性与有效性。