Visual to auditory sensory substitution devices convert visual information into sound and can provide valuable assistance for blind people. Recent iterations of these devices rely on depth sensors. Rules for converting depth into sound (i.e. the sonifications) are often designed arbitrarily, with no strong evidence for choosing one over another. The purpose of this work is to compare and understand the effectiveness of five depth sonifications in order to assist the design process of future visual to auditory systems for blind people which rely on depth sensors. The frequency, amplitude and reverberation of the sound as well as the repetition rate of short high-pitched sounds and the signal-to-noise ratio of a mixture between pure sound and noise are studied. We conducted positioning experiments with twenty-eight sighted blindfolded participants. Stage 1 incorporates learning phases followed by depth estimation tasks. Stage 2 adds the additional challenge of azimuth estimation to the first stage's protocol. Stage 3 tests learning retention by incorporating a 10-minute break before re-testing depth estimation. The best depth estimates in stage 1 were obtained with the sound frequency and the repetition rate of beeps. In stage 2, the beep repetition rate yielded the best depth estimation and no significant difference was observed for the azimuth estimation. Results of stage 3 showed that the beep repetition rate was the easiest sonification to memorize. Based on statistical analysis of the results, we discuss the effectiveness of each sonification and compare with other studies that encode depth into sounds. Finally we provide recommendations for the design of depth encoding.
翻译:视觉到听觉的感觉替代设备将视觉信息转换为声音,可为盲人提供宝贵帮助。这类设备的最新版本依赖于深度传感器。将深度转换为声音的规则(即听觉化)通常设计得较为随意,缺乏选择某一方式的有力依据。本研究旨在比较并理解五种深度听觉化的有效性,以辅助未来基于深度传感器的盲人视觉-听觉系统的设计过程。研究考察了声音的频率、振幅、混响、短促高音信号的重复率以及纯音与噪声混合的信噪比。我们招募了28名视力正常的蒙眼参与者进行定位实验。第一阶段包含学习阶段及随后的深度估计任务;第二阶段在第一阶段方案基础上增加了方位角估计的挑战;第三阶段在重新测试深度估计前加入10分钟休息,以检验学习保持效果。第一阶段中,声音频率和短促高音重复率获得了最佳的深度估计结果。第二阶段中,短促高音重复率在深度估计上表现最佳,而方位角估计未观察到显著差异。第三阶段结果表明,短促高音重复率是最易记忆的听觉化方式。基于对结果的统计分析,我们讨论了每种听觉化的有效性,并与将深度编码为声音的其他研究进行了比较,最后为深度编码设计提供了建议。