We describe a comprehensive methodology for developing user-voice personalized automatic speech recognition (ASR) models by effectively training models on mobile phones, allowing user data and models to be stored and used locally. To achieve this, we propose a resource-aware sub-model-based training approach that considers the RAM, and battery capabilities of mobile phones. By considering the evaluation metric and resource constraints of the mobile phones, we are able to perform efficient training and halt the process accordingly. To simulate real users, we use speakers with various accents. The entire on-device training and evaluation framework was then tested on various mobile phones across brands. We show that fine-tuning the models and selecting the right hyperparameter values is a trade-off between the lowest achievable performance metric, on-device training time, and memory consumption. Overall, our methodology offers a comprehensive solution for developing personalized ASR models while leveraging the capabilities of mobile phones, and balancing the need for accuracy with resource constraints.
翻译:本文描述了一种通过在手机上高效训练模型来实现用户语音个性化自动语音识别(ASR)模型的完整方法,从而允许用户数据和模型在本地存储和使用。为此,我们提出了一种基于子模型、考虑手机RAM和电池能力的资源感知训练方法。通过结合评估指标与手机的资源约束,我们能够高效执行训练并根据需要适时终止。为模拟真实用户,我们使用了多口音发音人。随后,我们在不同品牌的多种手机上测试了完整的端侧训练与评估框架。研究表明,模型微调与选择合适超参数的过程,需要在可获得的最低性能指标、端侧训练时间与内存消耗之间进行权衡。总体而言,本文提供了一种综合解决方案,在利用手机能力的同时,平衡了精度需求与资源约束,从而开发个性化的ASR模型。