Murine rodents generate ultrasonic vocalizations (USVs) with frequencies that extend to around 120kHz. These calls are important in social behaviour, and so their analysis can provide insights into the function of vocal communication, and its dysfunction. The manual identification of USVs, and subsequent classification into different subcategories is time consuming. Although machine learning approaches for identification and classification can lead to enormous efficiency gains, the time and effort required to generate training data can be high, and the accuracy of current approaches can be problematic. Here we compare the detection and classification performance of a trained human against two convolutional neural networks (CNNs), DeepSqueak and VocalMat, on audio containing rat USVs. Furthermore, we test the effect of inserting synthetic USVs into the training data of the VocalMat CNN as a means of reducing the workload associated with generating a training set. Our results indicate that VocalMat outperformed the DeepSqueak CNN on measures of call identification, and classification. Additionally, we found that the augmentation of training data with synthetic images resulted in a further improvement in accuracy, such that it was sufficiently close to human performance to allow for the use of this software in laboratory conditions.
翻译:啮齿类动物会产生频率延伸至约120kHz的超声波发声。这些叫声在社交行为中具有重要意义,因此对其进行分析有助于揭示声音交流的功能及其功能障碍。人工识别超声波发声并随后将其分类为不同子类别非常耗时。尽管用于识别和分类的机器学习方法能够带来巨大的效率提升,但生成训练数据所需的时间和精力可能很高,且当前方法的准确性可能存在一定问题。本研究将经过训练的人类与两种卷积神经网络(DeepSqueak和VocalMat)在含有大鼠超声波发声的音频上的检测和分类性能进行了比较。此外,我们还测试了将合成超声波发声图像插入VocalMat卷积神经网络训练数据中的效果,以此作为减少生成训练集工作量的手段。结果表明,在叫声识别和分类指标上,VocalMat的性能优于DeepSqueak卷积神经网络。此外,我们发现在训练数据中加入合成图像后,准确率进一步提升,以至于该软件的准确率与人类表现足够接近,使其能够在实验室条件下投入使用。