As 3D Gaussian Splatting (3DGS) emerges as a leading approach for novel view synthesis and scene reconstruction, its potential in digital asset creation has gained significant attention. An increasing number of asset libraries based on GS are being established. However, generating physics-based dynamic assets remains a time-consuming and expertise-intensive task, especially for non-experts. In this paper, we propose LIVE-GS, a highly realistic Virtual Reality (VR) system powered by Large Language Models (LLMs), which enables rapid creation of dynamic Gaussian assets and real-time VR interactions. To inform our system design, we conducted interviews to examine challenges faced by current GS-based VR systems and the specific demands of users. Based on these insights, we employed GPT-4o to analyze key physical properties of objects that significantly impact user interactions, ensuring physics-based interactions in VR align with real-world phenomena. A key innovation of LIVE-GS is its ability to predict reasonable parameters in just 10 seconds from static Gaussian assets while maintaining high-quality VR interactions. To validate our approach, we invited participants experienced in physical simulation to manually adjust physical parameters, providing a baseline for comparison in both asset quality and authoring efficiency. We also conducted a comprehensive user study to evaluate system usability and user satisfaction. Experimental results demonstrate that LIVE-GS, leveraging LLMs' scene understanding capabilities, can achieve efficient physical scene creation and natural interactions without requiring manual design or annotation.
翻译:随着三维高斯泼溅(3DGS)成为新视角合成与场景重建的主流方法,其在数字资产创建中的潜力日益受到关注。基于GS的资产库正在加速构建,但生成支持物理动态的资产仍是一项耗时且依赖专业知识的任务,对非专家用户尤为困难。本文提出LIVE-GS——一种由大语言模型(LLM)驱动的高逼真度虚拟现实(VR)系统,可实现动态高斯资产的快速创建与实时VR交互。为优化系统设计,我们通过用户访谈剖析了现有GS-VR系统的挑战及用户具体需求。基于这些洞察,我们利用GPT-4o分析对用户交互至关重要的物体物理属性,确保VR中的物理交互与真实世界现象一致。LIVE-GS的核心创新在于:仅需10秒即可从静态高斯资产中预测合理物理参数,同时保持高质量的VR交互。为验证方法有效性,我们邀请具有物理仿真经验的参与者手动调整物理参数,作为资产质量与创作效率的对比基准,并开展了综合用户研究评估系统可用性与满意度。实验结果表明,LIVE-GS借助LLM的场景理解能力,可在无需人工设计或标注的情况下实现高效的物理场景创建与自然交互。