Intuitively editing the appearance of materials from a single image is a challenging task given the complexity of the interactions between light and matter, and the ambivalence of human perception. This problem has been traditionally addressed by estimating additional factors of the scene like geometry or illumination, thus solving an inverse rendering problem and subduing the final quality of the results to the quality of these estimations. We present a single-image appearance editing framework that allows us to intuitively modify the material appearance of an object by increasing or decreasing high-level perceptual attributes describing such appearance (e.g., glossy or metallic). Our framework takes as input an in-the-wild image of a single object, where geometry, material, and illumination are not controlled, and inverse rendering is not required. We rely on generative models and devise a novel architecture with Selective Transfer Unit (STU) cells that allow to preserve the high-frequency details from the input image in the edited one. To train our framework we leverage a dataset with pairs of synthetic images rendered with physically-based algorithms, and the corresponding crowd-sourced ratings of high-level perceptual attributes. We show that our material editing framework outperforms the state of the art, and showcase its applicability on synthetic images, in-the-wild real-world photographs, and video sequences.
翻译:[translated abstract in Chinese]
从单张图像直观编辑材质外观是一项具有挑战性的任务,这源于光与物质相互作用的复杂性以及人类感知的歧义性。传统方法通过估计场景中的几何或光照等附加因素来应对该问题,即求解逆渲染问题,但最终结果的质量受制于这些估计的精度。本文提出了一种单图像外观编辑框架,允许通过增减描述材质外观的高层感知属性(如光泽度或金属感)来直观地修改物体材质外观。该框架以单张野外拍摄的物体图像为输入,无需控制几何、材质和光照条件,也无需进行逆渲染。我们基于生成模型设计了一种新颖的架构,其中包含选择性传输单元(STU)模块,可在编辑后的图像中保留输入图像的高频细节。为训练该框架,我们利用了一组配对数据:采用物理基算法渲染的合成图像及其对应的众包标注的高层感知属性评分。实验表明,我们的材质编辑框架优于现有技术,并在合成图像、野外真实照片及视频序列中验证了其实用性。