Recent advances in large language models (LLMs) have significantly improved language-driven 3D content generation, but most existing approaches still treat scene generation and user interaction as separate processes, limiting the adaptability and immersive potential of interactive multimedia systems. This paper presents a unified framework that closes the loop between language-driven 3D scene generation and immersive user interaction. Given natural language instructions, the system first constructs structured scene representations using LLMs, and then optimizes spatial layouts via reinforcement learning under geometric and semantic constraints. The generated environments are deployed in a virtual reality setting to facilitate HRI-in-the-loop, where user interactions provide continuous feedback to align generated content with human perception and usability. By tightly coupling generation and interaction, the proposed framework enables more responsive, adaptive, and realistic multimedia experiences. Experiments on the ALFRED benchmark demonstrate state-of-the-art performance in task-based scene generation. Furthermore, qualitative results and user studies show consistent improvements in immersion, interaction quality, and task efficiency, highlighting the importance of closed-loop integration of generation and interaction for next-generation multimedia systems. Our project page can be found at https://proj-showcase.github.io/h3ds/.
翻译:近期大语言模型的进展显著提升了语言驱动的三维内容生成能力,但现有方法仍将场景生成与用户交互视为独立过程,制约了交互式多媒体系统的适应性与沉浸潜力。本文提出一种统一框架,通过闭环架构实现语言驱动三维场景生成与沉浸式用户交互的深度融合。系统首先基于自然语言指令,利用大语言模型构建结构化场景表征,随后在几何与语义约束下通过强化学习优化空间布局。所生成的场景部署于虚拟现实环境,形成人机交互闭环——用户交互提供持续反馈,使生成内容与人类感知及可用性保持对齐。通过紧密耦合生成与交互环节,该框架能实现更敏捷、自适应且逼真的多媒体体验。在ALFRED基准上的实验表明,该方法在任务导向型场景生成中达到最先进性能。此外,定性结果与用户研究显示,该方法在沉浸感、交互质量与任务效率方面均有持续提升,凸显了生成-交互闭环集成对下一代多媒体系统的关键价值。项目页面详见https://proj-showcase.github.io/h3ds/。