Leader-follower interaction is an important paradigm in human-robot interaction (HRI). Yet, assigning roles in real time remains challenging for resource-constrained mobile and assistive robots. While large language models (LLMs) have shown promise for natural communication, their size and latency limit on-device deployment. Small language models (SLMs) offer a potential alternative, but their effectiveness for role classification in HRI has not been systematically evaluated. In this paper, we present a benchmark of SLMs for leader-follower communication, introducing a novel dataset derived from a published database and augmented with synthetic samples to capture interaction-specific dynamics. We investigate two adaptation strategies: prompt engineering and fine-tuning, studied under zero-shot and one-shot interaction modes, compared with an untrained baseline. Experiments with Qwen2.5-0.5B reveal that zero-shot fine-tuning achieves robust classification performance (86.66% accuracy) while maintaining low latency (22.2 ms per sample), significantly outperforming baseline and prompt-engineered approaches. However, results also indicate a performance degradation in one-shot modes, where increased context length challenges the model's architectural capacity. These findings demonstrate that fine-tuned SLMs provide an effective solution for direct role assignment, while highlighting critical trade-offs between dialogue complexity and classification reliability on the edge.
翻译:领导者-追随者交互是人机交互(HRI)中的重要范式。然而,对于资源受限的移动机器人和辅助机器人而言,实时分配角色仍具挑战性。尽管大型语言模型(LLMs)在自然语言交互中展现出潜力,但其规模大、延迟高限制了设备端部署。小语言模型(SLMs)作为一种潜在替代方案,其在HRI角色分类任务中的有效性尚未得到系统评估。本文提出一个面向领导者-追随者交互的SLMs基准测试,引入了一个基于已发表数据库构建的新数据集,并通过合成样本增强来捕捉交互特有动态。我们研究了两种适应策略:提示工程与微调,分别在零样本和单样本交互模式下进行实验,并与未训练基线进行比较。基于Qwen2.5-0.5B模型的实验表明,零样本微调在实现稳健分类性能(准确率86.66%)的同时保持低延迟(每样本22.2毫秒),显著优于基线方法和提示工程方法。然而,结果也显示在单样本模式下性能下降,此时增加的上下文长度挑战了模型架构容量。这些发现证明微调后的SLMs为直接角色分配提供了有效解决方案,同时揭示了边缘计算中对话复杂度与分类可靠性之间的关键权衡。