This work introduces a modular platform that brings together six AI services, automatic speech recognition via OpenAI Whisper, multilingual translation through Meta NLLB, speech synthesis using AWS Polly, emotion classification with RoBERTa, dialogue summarisation via flan t5 base samsum, and International Sign (IS) rendering through Google MediaPipe. A corpus of IS gesture recordings was processed to derive hand landmark coordinates, which were subsequently mapped onto three dimensional avatar animations inside a virtual reality (VR) environment. Validation comprised technical benchmarking of each AI component, including comparative assessments of speech synthesis providers and multilingual translation models (NLLB 200 and EuroLLM 1.7B variants). Technical evaluations confirmed the suitability of the platform for real time XR deployment. Speech synthesis benchmarking established that AWS Polly delivers the lowest latency at a competitive price point. The EuroLLM 1.7B Instruct variant attained a higher BLEU score, surpassing NLLB. These findings establish the viability of orchestrating cross modal AI services within XR settings for accessible, multilingual language instruction. The modular design permits independent scaling and adaptation to varied educational contexts, providing a foundation for equitable learning solutions aligned with European Union digital accessibility goals.
翻译:本文介绍了一个模块化平台,该平台整合了六项人工智能服务:基于OpenAI Whisper的自动语音识别、Meta NLLB的多语言翻译、AWS Polly的语音合成、RoBERTa的情感分类、flan t5 base samsum的对话摘要,以及通过Google MediaPipe实现的国际手语渲染。研究团队处理了国际手语手势记录语料库,提取手部地标坐标,并将其映射至虚拟现实环境中的三维虚拟角色动画中。验证工作包含各AI组件的技术基准测试,涵盖语音合成服务提供商及多语言翻译模型(NLLB 200与EuroLLM 1.7B变体)的对比评估。技术评估证实该平台适用于实时扩展现实部署。语音合成基准测试表明,AWS Polly在保持竞争性价格的同时实现了最低延迟;EuroLLM 1.7B指令微调变体则获得更高BLEU分数,优于NLLB。这些发现验证了在扩展现实环境中编排跨模态AI服务以实现无障碍多语教学的可行性。模块化设计允许各组件独立扩展并适配不同教育场景,为符合欧盟数字无障碍目标的公平教育解决方案奠定了基础。