Voice plays an important role in our lives by facilitating communication, conveying emotions, and indicating health. Therefore, tracking vocal interactions can provide valuable insight into many aspects of our lives. This paper presents our ongoing efforts to design a new vocal tracking system we call VoCopilot. VoCopilot is an end-to-end system centered around an energy-efficient acoustic hardware and firmware combined with advanced machine learning models. As a result, VoCopilot is able to continuously track conversations, record them, transcribe them, and then extract useful insights from them. By utilizing large language models, VoCopilot ensures the user can extract useful insights from recorded interactions without having to learn complex machine learning techniques. In order to protect the privacy of end users, VoCopilot uses a novel wake-up mechanism that only records conversations of end users. Additionally, all the rest of pipeline can be run on a commodity computer (Mac Mini M2). In this work, we show the effectiveness of VoCopilot in real-world environment for two use cases.
翻译:声音在我们的生活中扮演着重要角色,它促进交流、传递情感并指示健康状况。因此,追踪语音交互能为我们生活的诸多方面提供宝贵洞察。本文介绍了我们正在进行的系统设计工作——VoCopilot,一种新型语音追踪系统。VoCopilot是一个端到端系统,以高能效声学硬件与固件为核心,结合先进机器学习模型。最终,VoCopilot能持续追踪对话、记录并转录对话内容,进而从中提取有用信息。通过利用大型语言模型,VoCopilot确保用户无需掌握复杂机器学习技术即可从记录的交互中提取有效信息。为保护终端用户隐私,VoCopilot采用新颖的唤醒机制,仅记录终端用户本人的对话。此外,整个处理流程均可运行于普通计算机(Mac Mini M2)上。本文通过两个实际用例展示了VoCopilot在真实环境中的有效性。