Facial behavior constitutes the primary medium of human nonverbal communication. Existing synthesis methods predominantly follow two paradigms: coarse emotion category labels or one-hot Action Unit (AU) vectors from the Facial Action Coding System (FACS). Neither paradigm reliably renders fine-grained facial behaviors nor resolves anatomically implausible artifacts caused by conflicting AUs. Therefore, we propose a novel task paradigm: anatomically grounded facial behavior synthesis from FACS-based AU descriptions. This paradigm explicitly encodes FACS-defined muscle movement rules, inter-AU interactions, and conflict resolution mechanisms into natural language control signals. To enable systematic research, we develop a dynamic AU text processor, a FACS rule-based module that converts raw AU annotations into anatomically consistent natural language descriptions. Using this processor, we construct BP4D-AUText, the first large-scale text-image paired dataset for fine-grained facial behavior synthesis, comprising over 302K high-quality samples. Given that existing general semantic consistency metrics cannot capture the alignment between anatomical facial descriptions and synthesized muscle movements, we propose the Alignment Accuracy of AU Probability Distributions (AAAD), a task-specific metric that quantifies semantic consistency. Finally, we design VQ-AUFace, a robust baseline framework incorporating anatomical priors and progressive cross-modal alignment, to validate the paradigm. Extensive quantitative experiments and user studies demonstrate the paradigm significantly outperforms state-of-the-art methods, particularly in challenging conflicting AU scenarios, achieving superior anatomical fidelity, semantic consistency, and visual quality.
翻译:面部行为是人类非语言交流的主要媒介。现有合成方法主要遵循两种范式:粗粒度情感类别标签或来自面部动作编码系统(FACS)的独热(one-hot)动作单元(AU)向量。这两种范式均无法可靠地生成细粒度面部行为,也无法解决由冲突AU导致的解剖学不合理伪影。为此,我们提出一种新任务范式:基于FACS的AU描述进行解剖学基础的面部行为合成。该范式将FACS定义的肌肉运动规则、AU间交互作用及冲突解决机制显式编码为自然语言控制信号。为实现系统性研究,我们开发了动态AU文本处理器——一个基于FACS规则的模块,可将原始AU标注转换为解剖学一致的自然语言描述。利用该处理器,我们构建了BP4D-AUText——首个用于细粒度面部行为合成的大规模文本-图像配对数据集,包含超过30.2万高质量样本。鉴于现有通用语义一致性指标无法捕捉解剖面部描述与合成肌肉运动之间的对齐程度,我们提出AU概率分布对齐准确度(AAAD)——一种量化语义一致性的任务专属指标。最后,我们设计了VQ-AUFace——一个融入解剖先验和渐进式跨模态对齐的稳健基线框架,以验证该范式。大量定量实验和用户研究表明,该范式显著优于现有最先进方法,尤其在处理冲突AU挑战性场景时,在解剖保真度、语义一致性和视觉质量方面均表现卓越。