This paper proposes Allophant, a multilingual phoneme recognizer. It requires only a phoneme inventory for cross-lingual transfer to a target language, allowing for low-resource recognition. The architecture combines a compositional phone embedding approach with individually supervised phonetic attribute classifiers in a multi-task architecture. We also introduce Allophoible, an extension of the PHOIBLE database. When combined with a distance based mapping approach for grapheme-to-phoneme outputs, it allows us to train on PHOIBLE inventories directly. By training and evaluating on 34 languages, we found that the addition of multi-task learning improves the model's capability of being applied to unseen phonemes and phoneme inventories. On supervised languages we achieve phoneme error rate improvements of 11 percentage points (pp.) compared to a baseline without multi-task learning. Evaluation of zero-shot transfer on 84 languages yielded a decrease in PER of 2.63 pp. over the baseline.
翻译:本文提出了一种多语言音素识别器Allophant。该方法仅需目标语言的音素清单即可实现跨语言迁移,从而支持低资源场景下的识别。其架构通过组合式音素嵌入方法与多任务架构中单独监督的发音属性分类器相结合。我们还引入了PHOIBLE数据库的扩展版本Allophoible。结合基于距离映射的字素到音素输出方法,该扩展版本可直接利用PHOIBLE音素清单进行训练。通过在34种语言上的训练与评估,我们发现多任务学习的加入提升了模型对未见音素及音素清单的泛化能力。在有监督语言上,与未采用多任务学习的基线相比,音素错误率降低了11个百分点。对84种语言的零样本迁移评估显示,其音素错误率相比基线下降了2.63个百分点。