High-quality HDRIs(High Dynamic Range Images), typically HDR panoramas, are one of the most popular ways to create photorealistic lighting and 360-degree reflections of 3D scenes in graphics. Given the difficulty of capturing HDRIs, a versatile and controllable generative model is highly desired, where layman users can intuitively control the generation process. However, existing state-of-the-art methods still struggle to synthesize high-quality panoramas for complex scenes. In this work, we propose a zero-shot text-driven framework, Text2Light, to generate 4K+ resolution HDRIs without paired training data. Given a free-form text as the description of the scene, we synthesize the corresponding HDRI with two dedicated steps: 1) text-driven panorama generation in low dynamic range(LDR) and low resolution, and 2) super-resolution inverse tone mapping to scale up the LDR panorama both in resolution and dynamic range. Specifically, to achieve zero-shot text-driven panorama generation, we first build dual codebooks as the discrete representation for diverse environmental textures. Then, driven by the pre-trained CLIP model, a text-conditioned global sampler learns to sample holistic semantics from the global codebook according to the input text. Furthermore, a structure-aware local sampler learns to synthesize LDR panoramas patch-by-patch, guided by holistic semantics. To achieve super-resolution inverse tone mapping, we derive a continuous representation of 360-degree imaging from the LDR panorama as a set of structured latent codes anchored to the sphere. This continuous representation enables a versatile module to upscale the resolution and dynamic range simultaneously. Extensive experiments demonstrate the superior capability of Text2Light in generating high-quality HDR panoramas. In addition, we show the feasibility of our work in realistic rendering and immersive VR.
翻译:高质量HDRI(高动态范围图像),特别是HDR全景图,是图形学中创建三维场景逼真光照与360度反射的最流行方式之一。鉴于HDRI获取的困难性,亟需一种通用且可控的生成模型,使普通用户能够直观地控制生成过程。然而,现有最先进方法仍难以合成复杂场景的高质量全景图。本文提出零样本文本驱动框架Text2Light,可在无配对训练数据的情况下生成4K+分辨率HDRI。给定场景描述的自由形式文本,我们通过两个专用步骤合成对应HDRI:1)低动态范围(LDR)与低分辨率下的文本驱动全景图生成;2)超分辨率逆色调映射,同时提升LDR全景图的分辨率与动态范围。具体而言,为实现零样本文本驱动全景图生成,我们首先构建双码本作为多样化环境纹理的离散表示。随后,基于预训练CLIP模型,文本条件全局采样器根据输入文本从全局码本中学习采样整体语义。此外,结构感知局部采样器在整体语义引导下逐块学习合成LDR全景图。为实现超分辨率逆色调映射,我们从LDR全景图中导出360度成像的连续表示,作为锚定于球面的结构化隐码集合。该连续表示使通用模块能够同时提升分辨率与动态范围。大量实验证明Text2Light在生成高质量HDR全景图方面的卓越能力。此外,我们展示了本工作在真实感渲染与沉浸式VR中的可行性。