On-device automatic speech recognition systems face several challenges compared to server-based systems. They have to meet stricter constraints in terms of speed, disk size and memory while maintaining the same accuracy. Often they have to serve several applications with different distributions at once, such as communicating with a virtual assistant and speech-to-text. The simplest solution to serve multiple applications is to build application-specific (language) models, but this leads to an increase in memory. Therefore, we explore different data- and architecture-driven language modeling approaches to build a single application-agnostic model. We propose two novel feed-forward architectures that find an optimal trade off between different on-device constraints. In comparison to the application-specific solution, one of our novel approaches reduces the disk size by half, while maintaining speed and accuracy of the original model.
翻译:设备端自动语音识别系统相比服务器端系统面临多项挑战。它们必须在保持同等准确率的同时,满足更严格的响应速度、存储空间和内存限制,且通常需同时服务于虚拟助手交互、语音转文字等分布特性各异的多个应用。为支持多应用场景,最简方案是构建应用专属语言模型,但这会导致内存消耗增加。为此,我们探索了多种数据驱动与架构驱动的语言建模方法,旨在构建单一的应用无关模型。我们提出两种新型前馈式架构,可在不同设备端约束条件间实现最优权衡。与专属应用方案相比,其中一种新架构在保持原始模型速度和准确率的前提下,将存储空间需求缩减一半。