This paper introduces a non-native speech corpus consisting of narratives from fifty 5- to 6-year-old Chinese-English children. Transcripts totaling 6.5 hours of children taking a narrative comprehension test in English (L2) are presented, along with human-rated scores and annotations of grammatical and pronunciation errors. The children also completed the parallel MAIN tests in Chinese (L1) for reference purposes. For all tests we recorded audio and video with our innovative self-developed remote collection methods. The video recordings serve to mitigate the challenge of low intelligibility in L2 narratives produced by young children during the transcription process. This corpus offers valuable resources for second language teaching and has the potential to enhance the overall performance of automatic speech recognition (ASR).
翻译:本文介绍了一个包含五十名5至6岁中英双语儿童叙述内容的非母语音语语料库。语料库呈现了儿童用英语(第二语言)完成叙述理解测试时总计6.5小时的录音文本,并附有人工评分结果及语法与发音错误的标注。儿童还完成了中文(第一语言)的对应MAIN测试作为参照。所有测试均采用我们自主研发的创新远程采集方法进行音频与视频录制。视频记录有效缓解了低龄儿童第二语言叙述转录过程中可理解性差的难题。该语料库为第二语言教学提供了宝贵资源,并有望提升自动语音识别系统的整体性能。