According to the Stimulus Organism Response (SOR) theory, all human behavioral reactions are stimulated by context, where people will process the received stimulus and produce an appropriate reaction. This implies that in a specific context for a given input stimulus, a person can react differently according to their internal state and other contextual factors. Analogously, in dyadic interactions, humans communicate using verbal and nonverbal cues, where a broad spectrum of listeners' non-verbal reactions might be appropriate for responding to a specific speaker behaviour. There already exists a body of work that investigated the problem of automatically generating an appropriate reaction for a given input. However, none attempted to automatically generate multiple appropriate reactions in the context of dyadic interactions and evaluate the appropriateness of those reactions using objective measures. This paper starts by defining the facial Multiple Appropriate Reaction Generation (fMARG) task for the first time in the literature and proposes a new set of objective evaluation metrics to evaluate the appropriateness of the generated reactions. The paper subsequently introduces a framework to predict, generate, and evaluate multiple appropriate facial reactions.
翻译:根据刺激-机体-反应(SOR)理论,人类的所有行为反应均由情境激发,个体在接收到刺激后会对其进行加工并产生恰当的反应。这意味着在特定情境中,面对给定的输入刺激,个体可能根据其内部状态及其他情境因素作出不同反应。类似地,在双人交互中,人类通过言语与非言语线索进行交流,而听众的多种非言语反应都可能成为对特定说话者行为的恰当回应。已有研究探讨了针对给定输入自动生成恰当反应的问题,但尚无研究尝试在双人交互情境中自动生成多种恰当反应,并通过客观指标评估这些反应的恰当性。本文首次在文献中定义了面部多种恰当反应生成(fMARG)任务,并提出了一套新的客观评估指标以衡量生成反应的恰当性。随后,本文介绍了一个用于预测、生成及评估多种恰当面部反应的框架。