Intuitively editing the appearance of materials from a single image is a challenging task given the complexity of the interactions between light and matter, and the ambivalence of human perception. This problem has been traditionally addressed by estimating additional factors of the scene like geometry or illumination, thus solving an inverse rendering problem and subduing the final quality of the results to the quality of these estimations. We present a single-image appearance editing framework that allows us to intuitively modify the material appearance of an object by increasing or decreasing high-level perceptual attributes describing such appearance (e.g., glossy or metallic). Our framework takes as input an in-the-wild image of a single object, where geometry, material, and illumination are not controlled, and inverse rendering is not required. We rely on generative models and devise a novel architecture with Selective Transfer Unit (STU) cells that allow to preserve the high-frequency details from the input image in the edited one. To train our framework we leverage a dataset with pairs of synthetic images rendered with physically-based algorithms, and the corresponding crowd-sourced ratings of high-level perceptual attributes. We show that our material editing framework outperforms the state of the art, and showcase its applicability on synthetic images, in-the-wild real-world photographs, and video sequences.
翻译:[translated abstract in Chinese]
从单张图像中直观编辑材质外观是一项具有挑战性的任务,原因在于光与物质相互作用的复杂性以及人类感知的歧义性。传统方法通过估计场景的几何、光照等额外因素来解决这一问题,即求解逆渲染问题,但最终结果质量受限于这些估计的精度。我们提出了一种单图像外观编辑框架,通过增加或减少描述外观的高层感知属性(例如光泽度或金属感),能够直观地修改对象的材质外观。该框架以野外环境下单个物体的图像作为输入,无需控制几何、材质与光照条件,也无需进行逆渲染。我们借助生成模型,设计了一种包含选择性传输单元(STU)的新型网络架构,可在编辑后的图像中保留输入图像的高频细节。为训练该框架,我们使用了一组由物理算法渲染的合成图像对,并辅以相应的高层感知属性众包评分数据。实验表明,我们的材质编辑框架优于现有技术水平,并在合成图像、野外真实照片及视频序列中验证了其适用性。